Summary
Implement a new process scheduler inspired by the existing ULE scheduler, optimized for hybrid CPU topologies on recent Intel and AMD processors.
Motivation
Contemporary consumer CPUs from Intel and AMD feature a mix of core types with distinct performance characteristics (e.g., P/E/LP cores in Intel, x3D cache-enabled CCDs in AMD), making traditional symmetric core scheduling suboptimal. Adapting core priorities for process assignment can improve responsiveness and power efficiency, especially on laptops and desktop systems with hybrid architectures.
Technical Context
The existing ULE scheduler (sys/kern/sched_ule.c) provides a solid foundation for this work:
Key ULE Components to Build Upon:
CPU Search & Core Selection
sched_pickcpu() (line 1317-1449): Main CPU selection logic for thread placement
cpu_search_lowest() and cpu_search_highest() (lines 672-805): Tree-based CPU topology search
struct cpu_search (line 651-658): Encapsulates CPU search constraints and preferences
- Uses
cs_prefer to bias selection toward preferred cores; can be extended for hybrid core types
Load Balancing & Thread Movement
sched_balance() and sched_balance_pair() (lines 910-982): Long-term load rebalancing between CPUs
tdq_move() (lines 991-1015): Thread migration between queues; can be enhanced to detect CCD boundaries on AMD
tdq_steal() and runq_steal() (lines 1259-1271): Thread stealing for idle CPUs
Per-CPU Queue Management
struct tdq (line 238-265): Per-CPU thread dispatch queue with load tracking
tdq_load_add() / tdq_load_rem() (lines 555-587): Load accounting; can track per-core-class metrics
tdq_lowpri and priority tracking for preemption decisions
Thread Scheduling State
struct td_sched (line 90-104): Thread-specific scheduling metadata
ts_cpu: Can be extended to track core class (P/E/LP or x3D/non-x3D)
ts_flags: Room for new flags to indicate core class affinity requirements
Topology Discovery
struct cpu_group (referenced at line 245): CPU grouping hierarchy (via smp_topo_find() at line 1515)
- Topology already distinguishes cache levels and SMT; can be extended for core class detection
Sysctl Integration
- ULE registers multiple sysctl tunables (lines 3218-3259):
kern.sched.* namespace
- Can add new knobs:
kern.sched.hybrid_mode (favor P/E for Intel), kern.sched.amd_x3d_prefer, etc.
Requirements
Intel CPUs (Hybrid: P/E/LP cores)
- Schedule cores in this order of preference:
- P core (Performance core)
- E core (Efficiency core)
- Hyperthreaded sibling thread on P cores
- LP core (Low power core, e.g., found on some CPUs)
- Implement a sysctl tunable to favor P cores or E cores for top priority (other priorities remain as above).
- Use case: toggle for laptops to prioritize battery life (favor E cores) or performance (favor P cores).
- When number of runnable processes < total cores, prioritize actual P cores for initial placement.
- Prefer E cores or LP cores for new/burst processes if system is loaded (more processes than cores), but long-running tasks keep their placement preference.
- Do not starve any process regardless of core class.
AMD CPUs (x3D cache)
- For Ryzen 7950x3d, 7900x3d, 9950x3d, 9900x3d:
- Prioritize CCDs/cores with extra L3 cache ("x3d" cache) for process assignment.
- For less loaded systems, put most processes on x3d cache cores.
- If overloaded, new/burst processes go to non-x3d CCDs/cores to balance load.
- Once a process is on a CCD, avoid cross-CCD migration to reduce cache loss.
- For other AMD CPUs (no x3d): continue current ULE-like behavior applying SMT (hyperthreading) rules.
Core Algorithm Changes Needed:
-
CPU Topology Detection (via CPUID)
- Detect Intel hybrid core topology (P/E/LP distinction)
- Detect AMD x3D chipset topology from CPU model
- Mark cores with attributes in
struct cpu_group flags
-
Modified sched_pickcpu() Logic (refactor lines 1317-1449)
- Compute a "core class weight" for each candidate CPU
- Weight = function of (core_type, load, preferred_type_sysctl)
- Prefer higher-weight cores for new thread placement
-
Enhanced cpu_search_lowest() / cpu_search_highest()
- Accept a
core_class_preference parameter
- Score candidate CPUs by: (1) core class alignment, (2) load, (3) topology proximity
-
CCD Affinity for AMD (enhance tdq_move())
- Track which CCD a thread resides on
- Penalize (or prevent) cross-CCD migrations in load balancing
- Use lazy migration if beneficial for performance
-
New Thread Placement Strategy
- Detect if process is "new" (low runtime) vs "long-running"
- Route new/burst tasks to E cores / non-x3D cores when overloaded
- Keep long-running tasks on their preferred core class
-
Starvation Prevention
- Ensure all runnable threads eventually get scheduled
- Use periodic "fairness sweeps" to rotate threads between core classes
- Track per-core-class wait times; boost priority if waiting too long
Acceptance Criteria
- New scheduler is enabled as an experimental alternative to current ULE (e.g., via kernel config option
SCHED_HYBRID)
- Sysctl tunable for core preference documented and effective at runtime
- No starvation or excessive load on any core class during varied workloads
- Scheduler passes buildworld/kernel regression tests
- Documentation/man pages updated (man 4 cpu, man 8 sysctl for hybrid scheduler)
- Simple benchmark/test case included (e.g., CPU affinity verification)
Implementation Notes
- Start by forking
sys/kern/sched_ule.c to sys/kern/sched_hybrid.c or adding conditional compilation
- Register new scheduler as alternative via standard scheduler API (see
SCHED_ULE references)
- Update kernel build system to enable/disable hybrid scheduler per architecture (amd64 + arm64 initially)
- Consider backward compatibility: do not break existing ULE code path
AI-Assisted-by: Claude Sonnet
Summary
Implement a new process scheduler inspired by the existing ULE scheduler, optimized for hybrid CPU topologies on recent Intel and AMD processors.
Motivation
Contemporary consumer CPUs from Intel and AMD feature a mix of core types with distinct performance characteristics (e.g., P/E/LP cores in Intel, x3D cache-enabled CCDs in AMD), making traditional symmetric core scheduling suboptimal. Adapting core priorities for process assignment can improve responsiveness and power efficiency, especially on laptops and desktop systems with hybrid architectures.
Technical Context
The existing ULE scheduler (
sys/kern/sched_ule.c) provides a solid foundation for this work:Key ULE Components to Build Upon:
CPU Search & Core Selection
sched_pickcpu()(line 1317-1449): Main CPU selection logic for thread placementcpu_search_lowest()andcpu_search_highest()(lines 672-805): Tree-based CPU topology searchstruct cpu_search(line 651-658): Encapsulates CPU search constraints and preferencescs_preferto bias selection toward preferred cores; can be extended for hybrid core typesLoad Balancing & Thread Movement
sched_balance()andsched_balance_pair()(lines 910-982): Long-term load rebalancing between CPUstdq_move()(lines 991-1015): Thread migration between queues; can be enhanced to detect CCD boundaries on AMDtdq_steal()andrunq_steal()(lines 1259-1271): Thread stealing for idle CPUsPer-CPU Queue Management
struct tdq(line 238-265): Per-CPU thread dispatch queue with load trackingtdq_load_add()/tdq_load_rem()(lines 555-587): Load accounting; can track per-core-class metricstdq_lowpriand priority tracking for preemption decisionsThread Scheduling State
struct td_sched(line 90-104): Thread-specific scheduling metadatats_cpu: Can be extended to track core class (P/E/LP or x3D/non-x3D)ts_flags: Room for new flags to indicate core class affinity requirementsTopology Discovery
struct cpu_group(referenced at line 245): CPU grouping hierarchy (viasmp_topo_find()at line 1515)Sysctl Integration
kern.sched.*namespacekern.sched.hybrid_mode(favor P/E for Intel),kern.sched.amd_x3d_prefer, etc.Requirements
Intel CPUs (Hybrid: P/E/LP cores)
AMD CPUs (x3D cache)
Core Algorithm Changes Needed:
CPU Topology Detection (via
CPUID)struct cpu_groupflagsModified
sched_pickcpu()Logic (refactor lines 1317-1449)Enhanced
cpu_search_lowest()/cpu_search_highest()core_class_preferenceparameterCCD Affinity for AMD (enhance
tdq_move())New Thread Placement Strategy
Starvation Prevention
Acceptance Criteria
SCHED_HYBRID)Implementation Notes
sys/kern/sched_ule.ctosys/kern/sched_hybrid.cor adding conditional compilationSCHED_ULEreferences)AI-Assisted-by: Claude Sonnet