Skip to content

Add hybrid CPU-aware scheduler based on ULE for Intel P/E/LP and AMD x3D cores #346

Description

@laffer1

Summary

Implement a new process scheduler inspired by the existing ULE scheduler, optimized for hybrid CPU topologies on recent Intel and AMD processors.

Motivation

Contemporary consumer CPUs from Intel and AMD feature a mix of core types with distinct performance characteristics (e.g., P/E/LP cores in Intel, x3D cache-enabled CCDs in AMD), making traditional symmetric core scheduling suboptimal. Adapting core priorities for process assignment can improve responsiveness and power efficiency, especially on laptops and desktop systems with hybrid architectures.

Technical Context

The existing ULE scheduler (sys/kern/sched_ule.c) provides a solid foundation for this work:

Key ULE Components to Build Upon:

CPU Search & Core Selection

  • sched_pickcpu() (line 1317-1449): Main CPU selection logic for thread placement
  • cpu_search_lowest() and cpu_search_highest() (lines 672-805): Tree-based CPU topology search
  • struct cpu_search (line 651-658): Encapsulates CPU search constraints and preferences
  • Uses cs_prefer to bias selection toward preferred cores; can be extended for hybrid core types

Load Balancing & Thread Movement

  • sched_balance() and sched_balance_pair() (lines 910-982): Long-term load rebalancing between CPUs
  • tdq_move() (lines 991-1015): Thread migration between queues; can be enhanced to detect CCD boundaries on AMD
  • tdq_steal() and runq_steal() (lines 1259-1271): Thread stealing for idle CPUs

Per-CPU Queue Management

  • struct tdq (line 238-265): Per-CPU thread dispatch queue with load tracking
  • tdq_load_add() / tdq_load_rem() (lines 555-587): Load accounting; can track per-core-class metrics
  • tdq_lowpri and priority tracking for preemption decisions

Thread Scheduling State

  • struct td_sched (line 90-104): Thread-specific scheduling metadata
  • ts_cpu: Can be extended to track core class (P/E/LP or x3D/non-x3D)
  • ts_flags: Room for new flags to indicate core class affinity requirements

Topology Discovery

  • struct cpu_group (referenced at line 245): CPU grouping hierarchy (via smp_topo_find() at line 1515)
  • Topology already distinguishes cache levels and SMT; can be extended for core class detection

Sysctl Integration

  • ULE registers multiple sysctl tunables (lines 3218-3259): kern.sched.* namespace
  • Can add new knobs: kern.sched.hybrid_mode (favor P/E for Intel), kern.sched.amd_x3d_prefer, etc.

Requirements

Intel CPUs (Hybrid: P/E/LP cores)

  • Schedule cores in this order of preference:
    1. P core (Performance core)
    2. E core (Efficiency core)
    3. Hyperthreaded sibling thread on P cores
    4. LP core (Low power core, e.g., found on some CPUs)
  • Implement a sysctl tunable to favor P cores or E cores for top priority (other priorities remain as above).
    • Use case: toggle for laptops to prioritize battery life (favor E cores) or performance (favor P cores).
  • When number of runnable processes < total cores, prioritize actual P cores for initial placement.
  • Prefer E cores or LP cores for new/burst processes if system is loaded (more processes than cores), but long-running tasks keep their placement preference.
  • Do not starve any process regardless of core class.

AMD CPUs (x3D cache)

  • For Ryzen 7950x3d, 7900x3d, 9950x3d, 9900x3d:
    • Prioritize CCDs/cores with extra L3 cache ("x3d" cache) for process assignment.
    • For less loaded systems, put most processes on x3d cache cores.
    • If overloaded, new/burst processes go to non-x3d CCDs/cores to balance load.
    • Once a process is on a CCD, avoid cross-CCD migration to reduce cache loss.
  • For other AMD CPUs (no x3d): continue current ULE-like behavior applying SMT (hyperthreading) rules.

Core Algorithm Changes Needed:

  1. CPU Topology Detection (via CPUID)

    • Detect Intel hybrid core topology (P/E/LP distinction)
    • Detect AMD x3D chipset topology from CPU model
    • Mark cores with attributes in struct cpu_group flags
  2. Modified sched_pickcpu() Logic (refactor lines 1317-1449)

    • Compute a "core class weight" for each candidate CPU
    • Weight = function of (core_type, load, preferred_type_sysctl)
    • Prefer higher-weight cores for new thread placement
  3. Enhanced cpu_search_lowest() / cpu_search_highest()

    • Accept a core_class_preference parameter
    • Score candidate CPUs by: (1) core class alignment, (2) load, (3) topology proximity
  4. CCD Affinity for AMD (enhance tdq_move())

    • Track which CCD a thread resides on
    • Penalize (or prevent) cross-CCD migrations in load balancing
    • Use lazy migration if beneficial for performance
  5. New Thread Placement Strategy

    • Detect if process is "new" (low runtime) vs "long-running"
    • Route new/burst tasks to E cores / non-x3D cores when overloaded
    • Keep long-running tasks on their preferred core class
  6. Starvation Prevention

    • Ensure all runnable threads eventually get scheduled
    • Use periodic "fairness sweeps" to rotate threads between core classes
    • Track per-core-class wait times; boost priority if waiting too long

Acceptance Criteria

  • New scheduler is enabled as an experimental alternative to current ULE (e.g., via kernel config option SCHED_HYBRID)
  • Sysctl tunable for core preference documented and effective at runtime
  • No starvation or excessive load on any core class during varied workloads
  • Scheduler passes buildworld/kernel regression tests
  • Documentation/man pages updated (man 4 cpu, man 8 sysctl for hybrid scheduler)
  • Simple benchmark/test case included (e.g., CPU affinity verification)

Implementation Notes

  • Start by forking sys/kern/sched_ule.c to sys/kern/sched_hybrid.c or adding conditional compilation
  • Register new scheduler as alternative via standard scheduler API (see SCHED_ULE references)
  • Update kernel build system to enable/disable hybrid scheduler per architecture (amd64 + arm64 initially)
  • Consider backward compatibility: do not break existing ULE code path

AI-Assisted-by: Claude Sonnet

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions