ARTICLE · INTELLIGENCE

战地情报 · 详情页

来自尧图项目组的一线实战观察与深度解析

Linux thermal framework 核心机制解析:策略协商与热建模

Linux thermal framework 核心机制解析:策略协商与热建模 1. 为什么 thermal framework 不是“温度驱动模块”而是一套策略协商机制很多人第一次看到thermal framework这个名字下意识会把它理解成“Linux 内核里管温度的模块”——就像cpufreq管频率、cpuidle管空闲一样。但这种直觉恰恰是深入理解 thermal 的第一道坎。我刚接触这个子系统时在某款 ARM64 工控板上反复观察到CPU 温度已经飙到 92°C风扇全速转了三分钟/sys/class/thermal/thermal_zone0/temp持续跳变可系统负载一点没降top里几个计算密集型进程依然满核跑着。当时第一反应是“thermal 没生效驱动写错了”后来翻了三天代码才意识到thermal framework 本身根本不做任何主动降温动作它只负责把“热”这件事翻译成内核能听懂的语言并组织各方坐下来谈判——谁该让步、让多少、什么时候让全靠策略协商不是硬编码逻辑。这背后是 Linux 内核功耗管理哲学的根本性转变从“中心化控制”走向“分布式协商”。早期的散热方案比如某些 BSP 厂商在arch/arm/mach-xxx/下硬编码的温控逻辑把风扇调速、CPU 频率限制、甚至直接关核都写死在平台代码里结果就是换一块散热稍差的 PCB整套逻辑就失效加一个 GPU 负载原有策略立刻失衡连进/proc/sys/kernel/都找不到调节入口。而 thermal framework 的设计目标是让硬件能力描述thermal sensor cooling device、软件策略governor、用户空间干预sysfs 接口三者解耦。它不关心你用的是 NTC 热敏电阻还是 I2C 温度芯片也不规定风扇必须 PWM 控制——它只定义一套通用接口struct thermal_zone_device描述“热区”struct thermal_cooling_device描述“冷源”struct thermal_governor描述“谈判规则”。所有具体实现都通过注册回调函数挂载进来。举个生活化的类比thermal framework 就像一栋写字楼的中央能源管理系统。它不直接拧开或关掉每个办公室的空调阀门也不决定哪间会议室该关灯。它只做三件事1读取每层楼的温湿度传感器数据thermal zone2汇总所有可用的制冷设备VRM 散热器、SoC 内部 throttling、外部 PWM 风扇的能力参数cooling device3根据预设策略比如“节能优先”或“性能优先”协调各设备出力比例governor。真正的执行由每个冷却设备自己的驱动完成——风扇驱动收到set_cur_state(3)就调 PWM 占空比CPU 频率驱动收到set_cur_state(5)就触发cpufreq_set_policy()。framework 层只传递“需要降温”的信号强度不越俎代庖。这也是为什么你在dmesg里几乎看不到 thermal 相关的 ERROR 日志——它本身没有“失败”概念。当thermal_zone_device_update()被调用时它只是按顺序遍历所有注册的 cooling device调用其get_max_state()获取当前最大能力再调用set_cur_state()下发目标状态。如果某个 cooling device 的set_cur_state()回调返回错误比如风扇驱动发现 PWM channel 失效framework 会默默跳过它继续尝试下一个。这种“尽力而为”的设计保证了即使部分冷却设备异常整个热管理仍能降级运行而不是整套崩溃。提示不要在drivers/thermal/目录下找“温控算法”。真正的策略逻辑藏在drivers/thermal/gov_*子目录里比如gov_step_wise.c实现阶梯式降温gov_bang_bang.c实现开关式控制gov_power_allocator.c则基于功耗模型动态分配冷却资源。这些 governor 才是 thermal framework 的“大脑”而 framework 本身只是“神经系统”。2. thermal zone 的本质不是物理传感器而是热行为建模单元很多工程师拿到一块新板子第一反应是“找 thermal sensor 驱动”然后在dts里配tsadc { status okay; };以为这样就接入 thermal framework 了。但实际调试中常遇到cat /sys/class/thermal/thermal_zone0/temp返回-ENODEV或者数值恒为 0。这时翻遍 sensor 驱动代码也没问题最后发现根源在thermal_zone的建模方式上——thermal zone 不等于 sensor而是一个包含温度感知、热容特性、散热路径的抽象热模型。以常见的 SoC 平台为例thermal_zone0往往对应 CPU cluster但它内部可能聚合了多个物理 sensorCPU core 上的 die temperature sensor、GPU block 旁的 thermal diode、甚至 PCB 上的 NTC。framework 并不强制要求每个 sensor 对应一个 zone相反它鼓励将具有相似热惯性、共享散热路径的部件归入同一 zone。比如 Rockchip RK3588 的thermal-zonesdts 节点里cpu_thermal: cpu-thermal { polling-delay-passive 1000; /* 被动降温时每秒轮询一次 */ polling-delay-active 500; /* 主动降温时每500ms轮询一次 */ thermal-sensors tsadc 0, tsadc 1; /* 两个ADC通道 */ trips THERMAL_TRIP(100000, ACTIVE, 0) /* 100°C 触发主动降温 */ THERMAL_TRIP(110000, CRITICAL, 0) /* 110°C 触发紧急关机 */ ; #thermal-cells 2; };注意这里thermal-sensors是数组trips定义的是温度阈值事件而非单个 sensor 值。framework 在更新 zone 温度时会调用每个 sensor 的.get_temp()回调然后根据thermal_zone_params中的slope和offset参数进行线性校准很多 sensor 原始 ADC 值需转换为摄氏度最后取所有 sensor 的加权平均或最大值作为 zone 当前温度。这个过程在thermal_zone_get_temp()函数中完成关键代码段如下// drivers/thermal/thermal_core.c int thermal_zone_get_temp(struct thermal_zone_device *tz, int *temp) { int ret, i; long temp_val 0; long temp_max LONG_MIN; for (i 0; i tz-num_thermal_sensors; i) { ret tz-thermal_sensors[i]-ops-get_temp( tz-thermal_sensors[i], temp_val); if (ret) continue; // 校准temp_calibrated temp_raw * slope / 1000 offset temp_val mult_frac(temp_val, tz-tzp-slope, 1000) tz-tzp-offset; if (temp_val temp_max) temp_max temp_val; } *temp temp_max; return 0; }这意味着如果你的板子上有 3 个温度传感器但只在 dts 里声明了tsadc 0那么thermal_zone_get_temp()就只会读取第一个 sensor忽略其余两个——zone 温度永远无法反映真实热点。更隐蔽的问题是slope和offset参数某次我们调试一款国产 MCU发现temp值比红外热像仪实测低 15°C查到最后是offset -2000即 -20°C被误写成offset 200020°C导致所有读数整体偏高。这类校准参数通常来自 sensor datasheet 的 transfer function 表格必须逐项验证。另一个常被忽视的建模要素是polling-delay-*。很多人以为这只是“轮询间隔”其实它决定了 thermal framework 的响应粒度。polling-delay-passive用于 passive cooling如降低 CPU 频率polling-delay-active用于 active cooling如启动风扇。当 zone 温度超过 trip pointframework 会立即触发 cooling device但后续的温度跟踪仍依赖 polling。如果polling-delay-active设为 50005 秒而风扇启动后温度在 2 秒内就回落framework 可能根本来不及检测到变化继续维持高风速造成噪音和功耗浪费。实测经验对于桌面级 CPU建议polling-delay-active≤ 200ms嵌入式 SoC 可放宽至 500ms但必须配合thermal_zone_device_update()的手动触发比如在风扇驱动里检测到转速稳定后主动调用。注意thermal_zone的trips数组必须严格按温度升序排列。内核在thermal_zone_trip_update()中使用二分查找定位当前 trip如果顺序错乱会导致trip_id计算错误进而触发错误的 cooling level。这是 dts 编译时无法检查的隐性 bug只能靠dmesg | grep thermal观察 trip event 日志来排查。3. cooling device 的注册陷阱state 映射不是线性编号而是能力阶跃当你在dts中配置好 thermal zone下一步通常是注册 cooling device。常见做法是在drivers/thermal/下新建一个cooling_device.c实现struct thermal_cooling_device_ops然后调用thermal_cooling_device_register()。但实际部署时常出现“风扇不转”或“CPU 频率不降”的问题。深挖后发现根源往往不在驱动逻辑而在state的映射关系上——cooling device 的 state 不是简单的 0/1/2/3 编号而是代表冷却能力的离散阶跃值且不同 device 的 state 含义完全独立。以 PWM 风扇为例它的get_max_state()返回 10get_cur_state()返回当前占空比档位0~10set_cur_state()则根据档位设置 PWM duty cycle。但这里的 “10” 并不意味着“100% 风速”而是一个相对能力标尺。假设你的风扇驱动这样实现static int fan_set_cur_state(struct thermal_cooling_device *cdev, unsigned long state) { struct fan_cooling_device *fan cdev-devdata; u8 duty_cycle; switch (state) { case 0: duty_cycle 0; break; /* 停转 */ case 1: duty_cycle 20; break; /* 20% */ case 2: duty_cycle 40; break; /* 40% */ case 3: duty_cycle 60; break; /* 60% */ case 4: duty_cycle 80; break; /* 80% */ case 5: duty_cycle 100; break; /* 100% */ default: return -EINVAL; } return pwm_config(fan-pwm, duty_cycle * 255 / 100, 255); /* 假设 PWM range 0-255 */ }表面看没问题但thermal framework在调用set_cur_state()时传入的state值是由 governor 根据当前温度与 trip 的偏差计算得出的。比如gov_step_wise会这样计算目标 state// drivers/thermal/gov_step_wise.c target_state (tz-temperature - trip-temperature) * (cdev-max_state - cdev-min_state) / (trip-hysteresis ? trip-hysteresis : 1000); target_state clamp_t(unsigned long, target_state, cdev-min_state, cdev-max_state);这里target_state是一个浮点运算后的整数可能得到 7、8、9 等值。但你的fan_set_cur_state()只处理 0~5遇到state7就返回-EINVALframework 会记录thermal cooling device set state failed日志并跳过该 device。这就是为什么风扇“明明注册成功却不动”的根本原因——state 空间必须连续覆盖min_state到max_state的全部整数不能有空缺。更隐蔽的陷阱是min_state的设定。默认情况下thermal_cooling_device_register()将min_state设为 0但某些 cooling device如 CPU frequency scaling的最小有效 state 并非 0。例如cpufreq_cooling_register()注册的 device其min_state实际对应freq_table[0]的频率而freq_table可能从 500MHz 开始没有 0Hz。如果 governor 计算出target_state0cpufreq_set_cur_state()会尝试设置最低频点这本身没问题但如果freq_table为空或初始化失败target_state0就可能导致cpufreq驱动 panic。实操中我们曾遇到一款 i.MX8M Plus 板子cpufreq_cooling_device注册后min_state0,max_state3但freq_table只有 2 个有效条目1.2GHz 和 1.6GHz。当温度略超 tripgovernor 计算target_state1cpufreq驱动却因freq_table[1]为空而返回-ENODEV。解决方案不是改 governor而是确保freq_table初始化完整或在cpufreq_cooling_register()前显式调用cpufreq_frequency_table_cpuinfo()验证表有效性。另一个关键细节是state的原子性。set_cur_state()必须是原子操作不能在中间被中断打断。某次我们在 RT-Preempt 内核上调试发现风扇转速忽高忽低dmesg里频繁出现thermal cooling device state changed。最终定位到pwm_config()调用中包含了mutex_lock()而该 mutex 在中断上下文被持有导致set_cur_state()被阻塞。修复方法是将 PWM 配置移到 workqueue 中异步执行set_cur_state()只负责提交任务。提示验证 cooling device 是否正常工作的最简单方法是手动写 sysfsecho 5 /sys/class/thermal/cooling_device0/cur_state cat /sys/class/thermal/cooling_device0/cur_state # 应返回 5如果返回值与写入值不符说明get_cur_state()实现有问题如果写入后无物理响应检查set_cur_state()的返回值dmesg | grep cooling。4. governor 的选择逻辑不是“越激进越好”而是匹配热惯性特征thermal framework 提供了多种 governorstep_wise、bang_bang、power_allocator、user_space很多工程师习惯性选择step_wise认为它“渐进式降温更安全”。但实际项目中我们发现bang_bang在嵌入式设备上反而更稳定power_allocator在高性能服务器上才能发挥价值。这背后的核心逻辑是governor 的选择必须匹配被控对象的热惯性thermal inertia和冷却设备的响应延迟response latency。热惯性指温度变化的滞后性。CPU core 的热惯性很小——负载突增后die temperature 在毫秒级内就会上升而 PCB 的热惯性很大同样负载下PCB 温度可能需要数分钟才达到平衡。冷却设备的响应延迟也差异巨大PWM 风扇从 0% 到 100% 转速需 1~2 秒CPU frequency scaling 几乎瞬时完成而液冷泵的启停则需 5~10 秒。bang_banggovernor 的逻辑极简温度低于 trip 下限 → 关闭所有 cooling温度高于 trip 上限 → 全力 cooling。它适合热惯性大、响应延迟长的系统。比如某款工业网关SoC 封装在密闭金属壳内PCB 散热全靠自然对流。实测热时间常数 τ ≈ 120 秒温度变化 63% 所需时间。若用step_wise每次温度微超就小幅降频CPU 频率在 1.0GHz/1.2GHz/1.4GHz 间反复切换不仅无法有效降温因为 PCB 温度变化太慢还导致系统吞吐量抖动。换成bang_bang后设定trip_low70000,trip_high85000CPU 在 70°C 以下全速运行一旦超 85°C 就立刻锁频至 800MHz风扇全速120 秒后 PCB 温度缓慢回落系统进入稳定低功耗态。dmesg日志显示bang_bang的thermal_zone_trip_update()调用频率仅为step_wise的 1/20大幅降低内核调度开销。step_wise则适合热惯性小、需要精细调控的场景。桌面级 CPU 的 die temperature 响应极快step_wise能根据温度偏离 trip 的程度线性调整 cooling level避免bang_bang的剧烈震荡。但它的致命弱点是polling-delay敏感。当polling-delay-active500时step_wise每 500ms 计算一次 target state。如果温度在两次 polling 之间快速跨越多个 tripstep_wise会累积误差导致 cooling level 过冲。我们曾在一个 X86 服务器上观察到温度从 75°C 突升至 95°C跨越 3 个 tripstep_wise在 1 秒内将风扇从 30% 调至 100%但温度已开始回落结果风扇持续满转 30 秒噪音超标。解决方案是启用thermal_zone_device_update()的 event-driven 模式——在 sensor 驱动中当 ADC 值变化超过阈值时主动触发thermal_zone_device_update()绕过 polling 机制。power_allocator是最复杂的 governor它不直接控制 cooling device 的 state而是基于功耗模型power model分配 cooling budget。其核心思想是每个 cooling device 有一个power属性单位 mWgovernor 计算当前 zone 需要移除的总功耗P_needed然后按power比例分配给各 device。例如CPU frequency scaling 的power为 5000mW风扇的power为 2000mW则P_needed7000mW时CPU 分配 5000mW即降频至对应功耗点风扇分配 2000mW即调至对应风速。这要求所有 cooling device 必须实现.get_requested_power()和.state2power()回调且 power model 必须准确。某次我们在 NVIDIA Jetson Orin 上启用power_allocator发现 GPU 温度失控查到最后是gpu_cooling_device的state2power()表格未校准——表格中 state3 对应 15W但实测 GPU 在该 state 下功耗仅 8W导致 governor 低估了 GPU 的冷却能力过度依赖风扇而风扇又因响应延迟无法及时补足形成恶性循环。经验技巧在drivers/thermal/目录下gov_*文件的编译选项CONFIG_THERMAL_GOV_*默认只选STEP_WISE。若要启用BANG_BANG必须在 kernel config 中显式开启CONFIG_THERMAL_GOV_BANG_BANGy否则thermal_register_governor()会失败dmesg显示thermal governor bang-bang not registered。这是很多新手卡住的第一步。5. 用户空间干预的边界sysfs 不是万能钥匙而是策略协商的输入端口很多工程师认为thermal framework 的终极控制权在用户空间——只要写echo 1 /sys/class/thermal/cooling_device0/cur_state就能强制风扇启动。但实际项目中我们发现这种操作常被 framework 忽略cur_state文件读出来还是 0。这是因为sysfs 接口不是直接操控硬件的“后门”而是 thermal framework 策略协商流程中的一个输入变量其效果受 governor 策略和 cooling device 状态约束。cooling_deviceX/cur_state文件的 write 操作最终调用cooling_device_sysfs_state_store()该函数会调用cdev-ops-set_cur_state()但前提是cdev-updated为 true 且cdev-device已注册。更重要的是cur_state的修改会被 governor 在下次thermal_zone_trip_update()时覆盖。比如step_wisegovernor 在每轮 polling 中都会根据当前温度重新计算target_state然后调用thermal_cooling_device_set_cur_state()强制重置。因此手动写cur_state只能在两次 polling 间隙生效且仅当 governor 未主动干预时才持久。真正可靠的用户空间干预方式是通过thermal_zoneX/type和thermal_zoneX/mode文件。type文件显示 zone 类型如cpu-thermal不可写mode文件则控制整个 zone 的使能状态# 查看当前模式 cat /sys/class/thermal/thermal_zone0/mode # 输出 enabled # 临时禁用 thermal 管理慎用 echo disabled /sys/class/thermal/thermal_zone0/mode # 恢复 echo enabled /sys/class/thermal/thermal_zone0/mode当modedisabled时framework 停止调用thermal_zone_device_update()所有 cooling device 的 state 保持不变governor 不再计算。这是调试 thermal 问题的黄金开关——比如你想确认是否是 thermal 导致 CPU 降频就先echo disabled再跑压力测试看频率是否稳定。但要注意disabled不等于关闭硬件sensor 仍在读数cooling device 仍保持最后状态只是 framework 不再主动干预。另一个重要接口是thermal_zoneX/policy它指定当前使用的 governor。默认为step_wise可切换为bang_bang或user_space# 切换 governor需 kernel config 支持 echo bang-bang /sys/class/thermal/thermal_zone0/policy但policy的切换并非即时生效。thermal_set_governor()函数会先调用原 governor 的.throttle()回调如果有再初始化新 governor 的状态。对于power_allocator这还包括重新加载 power model 数据。因此切换后需等待 1~2 个 polling 周期新策略才完全生效。最灵活的干预方式是user_spacegovernor。启用后framework 完全放弃自动策略将cur_state的控制权完全交给用户空间程序# 启用 user_space 模式 echo user_space /sys/class/thermal/thermal_zone0/policy # 此时 cooling device 的 cur_state 可被任意写入且不会被覆盖 echo 5 /sys/class/thermal/cooling_device0/cur_state但这要求用户空间程序必须承担全部热管理逻辑。我们曾为某款车载信息娱乐系统开发过 custom thermal daemon它读取/sys/class/thermal/thermal_zone0/temp结合车速、空调状态等外部信号用 PID 算法计算目标风扇转速再写入cooling_device0/cur_state。关键点在于daemon 必须以高优先级运行chrt -f 50 ./thermal_daemon并确保写入频率不低于 framework 的 polling rate否则cur_state会被user_spacegovernor 的空闲逻辑重置为 0。警告不要在thermal_zoneX/trip_point_*_temp文件中随意修改温度阈值。这些值在 kernel boot 时由 dts 解析固定运行时修改仅影响当前 trip 的比较值但不会更新thermal_zone的内部 trip table。强行修改可能导致thermal_zone_trip_update()计算错误甚至 kernel panic。如需动态调整应通过thermal_zone_device_update()重新触发 trip 评估或在用户空间 daemon 中实现自适应 trip 逻辑。6. 调试 thermal 问题的四层排查链路从 dmesg 到 trace-cmd 的完整路径当 thermal 行为异常如温度飙升却不降频、风扇狂转却无效很多工程师习惯性grep thermal /proc/kmsg或dmesg | grep -i thermal但往往只见零星日志无法定位根因。经过数十个项目实战我们总结出一套四层递进式排查链路覆盖从内核启动到运行时的全生命周期6.1 第一层dmesg 启动日志的隐藏线索kernel boot 阶段的 thermal 初始化日志是黄金信息源但常被忽略。重点搜索以下关键词thermal zone: 确认 zone 是否成功注册。正常输出类似thermal zone0: binding with thermal sensor tsadc.0。若出现thermal zone0: no thermal sensor found说明 dts 中thermal-sensors路径错误或 sensor 驱动未加载。cooling device: 查看 cooling device 注册状态。registered as cooling_device0表示成功failed to register cooling device则需检查thermal_cooling_device_register()的返回值。governor:registered thermal governor step_wise表示 governor 加载成功若缺失检查CONFIG_THERMAL_GOV_*是否启用。特别注意thermal zone的polling delay日志thermal zone0: polling delay set to 1000 ms。如果此处显示0 ms说明 dts 中polling-delay-*未正确解析framework 会使用默认值通常为 0导致 polling 频率过高CPU 占用率飙升。6.2 第二层sysfs 文件树的状态快照在系统运行时采集完整的 thermal sysfs 状态比dmesg更直观# 生成 thermal 状态快照 mkdir /tmp/thermal_debug for f in /sys/class/thermal/thermal_zone*/{type,mode,policy,temp,trips/*_temp}; do if [ -f $f ]; then echo $f /tmp/thermal_debug/status.log cat $f 2/dev/null /tmp/thermal_debug/status.log fi done for f in /sys/class/thermal/cooling_device*/{type,cur_state,max_state}; do if [ -f $f ]; then echo $f /tmp/thermal_debug/status.log cat $f 2/dev/null /tmp/thermal_debug/status.log fi done关键检查点thermal_zoneX/temp是否可读若返回0或-ENODEV检查 sensor 驱动get_temp()实现。cooling_deviceX/cur_state是否随温度变化若恒为 0检查 governor 是否启用及set_cur_state()返回值。thermal_zoneX/trips/*/temp的数值是否符合预期比如trip_point_0_temp应为 7000070°C而非 70误写为摄氏度。6.3 第三层function graph trace 定位执行瓶颈当怀疑 governor 策略未执行或 cooling device 调用失败时启用 ftrace# 启用 thermal 相关 tracepoint echo 1 /sys/kernel/debug/tracing/events/thermal/thermal_zone_trip_update/enable echo 1 /sys/kernel/debug/tracing/events/thermal/thermal_cooling_device_state_update/enable # 设置 trace buffer 大小 echo 1048576 /sys/kernel/debug/tracing/buffer_size_kb # 开始 trace echo 1 /sys/kernel/debug/tracing/tracing_on # 运行测试如 stress-ng -c 4 stress-ng -c 4 -t 30s # 停止 trace echo 0 /sys/kernel/debug/tracing/tracing_on # 查看结果 cat /sys/kernel/debug/tracing/trace典型 trace 输出thermal_zone_trip_update: thermal_zone0: trip0, temp85000, hyst1000 thermal_cooling_device_state_update: cooling_device0: state3 - 5 thermal_cooling_device_state_update: cooling_device1: state0 - 2若thermal_zone_trip_update无输出说明 zone 未被更新若只有trip_update无state_update说明 governor 未调用set_cur_state()若state_update显示state0 - 0说明set_cur_state()返回错误。6.4 第四层perf probe 动态注入断点对于更深层的逻辑错误如step_wise计算 target_state 错误需用 perf probe 在 kernel 函数内设断点# 在 thermal_zone_trip_update() 函数内设 probe perf probe -a thermal_zone_trip_update:0 temp0(%rdi) trip8(%rdi) # 记录 perf data perf record -e probe:thermal_zone_trip_update -aR sleep 30 # 分析 perf scriptprobe 输出会显示每次调用时的temp和trip值可验证温度读数是否异常、trip index 是否越界。某次我们发现trip255远超num_trips3根源是thermal_zone的trips数组内存越界dts 中trips定义错误。最后提醒所有调试操作应在非生产环境进行。perf probe和ftrace会显著增加内核开销echo disabled会关闭热保护务必在测试完成后恢复enabled模式。真正的稳定性永远建立在对 thermal framework 架构的透彻理解之上而非临时 patch。
RELATED READING

延伸阅读

更多一线实战笔记与深度复盘,助您持续精进