TLDR:

  1. CPU2026 延续 SPEC CPU 的定位:测量 natively compiled C/C++/Fortran 程序在通用 CPU 上的性能和能耗,并要求结果可验证、可复现、可审计。
  2. 相比 CPU2017,CPU2026 扩大到 52 个 benchmark component,引入更多应用域、更多多 workload benchmark、更多前端受限 integer workload 和更多多线程 benchmark。
  3. 论文的主体是 benchmark engineering:从 CPUv8 search program 收集真实开源应用,经过 adaptation、workload selection、performance characterization、portability hardening 后形成正式 suite。
  4. 论文强调 portability 与 representativeness 的张力。现代 AI inference、crypto、media codec 等重要应用由于高度依赖 ISA-specific 优化或难以保证 equal work,很多没有进入最终 suite。
  5. CPU2026 的新方法学贡献是 Rolling Round-Robin Rate (RRR):一种标准化、确定性的异构 multiprogrammed workload 运行方式,用于补足传统 SPECrate homogeneous multi-copy 方法的不足。

背景定位:SPEC CPU 是测量工具

论文首先明确 SPEC CPU 的角色。SPEC CPU 系列从 CPU89 到 CPU2017 已经服务通用处理器评估三十多年,其价值在于提供公平、可复现、可比较的测量基准。作者特别强调 SPEC CPU 像一把 tape measure:它提供测量工具和正式结果提交规则,但不规定研究者必须只测某一种 compiler flag、thread count、OS 或系统配置。

CPU2026 的范围边界很明确:

  1. 覆盖 natively compiled C、C++、Fortran 程序。
  2. 覆盖 single-threaded 和 multi-threaded 场景。
  3. 不覆盖 Java、Python、Julia 等 managed runtime,因为 JIT 和 runtime behavior 会引入可移植性和运行方差问题。
  4. 不只测“跑得快”,还必须验证“跑得正确”。每个 benchmark 都要产生可检查的输出,并在不同系统上满足 golden reference。
  5. 主要服务两类用户:少量用户提交 official score;更多用户把 suite 当作研究 harness,用于 compiler、microarchitecture、system software 和 energy 分析。

这部分奠定了全文的主线:CPU2026 的目标是在真实性、可移植性、确定性和长期可用性之间建立可执行的工程规范。

CPU2026 的关键变化(相对 CPU2017)

变化 内容
Suite diversity 扩展应用域和微架构行为,更多 benchmark 使用多个 workload
Scale and resource component 数量增加;SPECspeed large multi-threaded footprint 从 16 GB 增至 64 GB;SPECrate 仍为每 copy 2 GB
Language standards C18、C++17、Fortran 2018;并行方式包括 OpenMP、std::thread、Fortran DO CONCURRENT
Analysis capability raw output 增加 coefficient of variation、quartile、min/max/average copy time、standard deviation 等统计
RRR 新增 exhibition 形式的 Rolling Round-Robin Rate,用于异构 multiprogrammed workload
Reporting category 公有云 bare-metal 可以提交 compliant score;新增 vendor-supported compiler 与 community-supported open-source compiler 区分

suite 规模如下:

SPECrate®2026 Integer SPECspeed®2026 Integer speed MT[A] Language KLOC files[B] KLOC hit[C] Application Domain ref
801.xz_s MT C++, C 53 5 Data compression [15]
706.stockfish_r C++ 13 5 Game (chess) - A/B search, deep learning neural network [16]
707.ntest_r 807.ntest_s MT C++ 16 5 Game (othello) - A/B search with heuristic eval function [17]
708.sqlite_r C 245 23 SQL compiler/interpreter and database [18]
710.omnetpp_r C++, C 224 19 Discrete event modeling - network and queuing simulations [19]
714.cpython_r C 747 56 Python interpreter [20]
817.flac_s MT C++, C 57 7 Lossless audio codec [21]
721.gcc_r 821.gcc_s MT C++, C 3,833 326 C language optimizing compiler [22]
723.llvm_r 823.llvm_s MT C++, C 3,167 123 C/C++ language optimizing compiler [23]
727.cppcheck_r 827.cppcheck_s MT C++ 287 111 Static analysis of C/C++ code [24]
729.abc_r 829.abc_s C++, C 989 37 Sequential logic synthesis and formal verification [25]
734.vpr_r 834.vpr_s C++, C 210 30 FPGA circuit place and route [26]
735.gem5_r 835.gem5_s C++, C 971 78 Computer architecture simulation model [27]
838.diamond_s MT C++, C 239 12 Bioinformatics - metagenomics and protein sequencing [28]
846.minizinc_s MT C++, C 372 33 Constraint programming (solvers: gecode and chuffed) [29]
750.sealcrypto_r C++, C 39 5 Security and privacy - Homomorphically Encrypted query [30]
753.ns3_r 853.ns3_s C++ 942 71 Discrete event network simulator for internet systems [31]
854.graph500_s MT C 10 1 Graph analytics [32]
777.zstd_r C 58 7 Data compression/decompression [33]
SPECrate®2026 Floating Point SPECspeed®2026 Floating Point speed MT[A] Language KLOC files[B] KLOC hit[C] Application Domain ref
800.pot3d_s MT Fortran 12 1 Solar physics: finite diff method, conjugate gradient solver [34]
803.sph_exa_s MT C++ 3 1 Astrophysics - Smoothed Particle Hydrodynamics (SPH) [35]
709.cactus_r 809.cactus_s MT C++, C 187 19 Astrophysics - relativity, finite difference, time integration [36]
811.tealeaf_s MT C 5 1 High energy physics [37]
816.nab_s MT C 26 2 Molecular modeling [38]
820.cloverleaf_s MT Fortran 10 1 Explicit hydrodynamics [39]
722.palm_r 822.palm_s MT Fortran 218 6 Atmospheric science [40]
731.astcenc_r C++ 43 8 Computer vision - Adaptive Scalable Texture Compression [41]
736.ocio_r C++ 183 13 Color management for visual effects and animation [42]
737.gmsh_r C++, C 721 35 Finite element mesh generation [43]
748.flightdm_r C++ 100 14 Flight dynamics models for aeronautics [44]
749.fotonik3d_r 849.fotonik3d_s MT Fortran 15 2 Computational Electromagnetics (CEM) [45]
857.namd_s MT C++ 9 2 Classical molecular dynamics simulation [46]
765.roms_r 865.roms_s MT Fortran 585 13 Regional ocean modeling [47]
766.femflow_r C++ 2,505 23 Fluid dynamics: high-order finite element method [48]
767.nest_r 867.nest_s MT C++ 208 17 Neuroscience simulator for spiking neural network models [49]
772.marian_r 872.marian_s MT C++ 219 15 Neural machine translation for written language [50]
782.lbm_r C 1 1 Computational fluid dynamics, Lattice Boltzmann Method [51]
881.neutron_s MT C 4 1 Physics simulation of neutron transport in nuclear reactors [52]

从体系结构研究角度看,最重要的变化有三点:

  1. CPU2026 更强调 frontend-bound integer workload。论文指出现代软件 code footprint 更大、控制流更复杂,因此需要覆盖 ITLB miss、instruction delivery pressure、branch misprediction 等前端瓶颈。
  2. CPU2026 加入多个 multithreaded integer benchmark,弥补 CPU2017 在这方面的缺口。
  3. CPU2026 不再只给 homogeneous rate 运行方式,而是通过 RRR 给 multiprogramming 研究提供标准化 workload construction。

Benchmark Development:从真实应用到正式 benchmark

论文第三部分描述 benchmark development pipeline。这个 pipeline 并不是线性的,而是候选收集、适配、workload 选择、性能表征、裁剪并行推进。

CPUv8 Search Program

CPUv8 search program 从 2020 年 2 月持续到 2023 年 3 月,面向开源社区、学术界和工业界征集候选。结果:

  1. 33 个候选进入考虑。
  2. 29 个完成初始 porting 和 workload definition。
  3. 24 个外部候选最终进入 CPU2026 suite。

作者强调这些候选多数来自活跃开源项目。SPEC 委员会与上游社区协作,不仅将应用改造成 benchmark,也向上游提交 bug fix、portability fix 和 performance improvement。这一点是论文反复强调的“symbiotic relationship”:SPEC 获得真实应用,开源项目获得跨系统测试和工程修复。

Adaptation:控制环境变量,保证 equal work

真实应用不能直接作为 SPEC benchmark。原因是它们可能读取系统时间、使用随机数、依赖环境变量、调用系统资源限制、包含平台专用优化、进行大量 I/O,或者在不同编译器/ISA 上走不同路径。adaptation 的目标是让 benchmark 在任何 compliant system 上执行相同数量的 user-space work,并在容差内产生相同输出。

主要处理包括:

  1. 消除非确定性:高熵随机源替换为 deterministic PRNG,如 std::mt19937;不稳定算法替换为有稳定定义的形式,如 std::stable_sort
  2. 最大化可移植性:移除 hand-coded assembly、compiler intrinsic、非标准语言扩展,替换为 portable C/C++/Fortran。
  3. 隔离运行环境:删除或替换 getenvsetenvgettimeofdaygetrusagesetrlimitsignaldlopen 等会引入外部差异的接口。
  4. 聚焦 user-space execution:尽量减少 system call,目标是至少 95% 执行时间在 benchmark 提供的 user-space code 中。
  5. 抑制 threading artifact:从多线程程序派生 SPECrate 单线程版本时,去除 lock/mutex 等不反映单线程性能的“threading tax”,同时为 SPECspeed 保留多线程 hook。
  6. 跨系统验证:在 x86、ARM、POWER、RISC-V 相关系统,以及 Linux、Windows、macOS 等 OS 上,用 GCC、LLVM 和多个 vendor compiler 在 -O2-O3、LTO、PGO 等设置下构建和验证。
  7. 法务审查:所有源代码和数据输入都需要许可证和来源清晰,确保 suite 可商业分发。

Integer 与 Floating-Point 的分类

SPEC CPU 传统上分为 INT 和 FP。过去常用动态 FP 指令比例判断:超过 10% 归入 FP,低于 1% 归入 INT。但现代 SIMD 执行资源共享、整数程序也可能使用浮点 load/store 指令搬运向量数据,导致 1% 到 10% 的 gray zone 增多。

CPU2026 对 gray zone 采用定性判断:根据应用的主要计算目的和用户社区认知分类

Benchmark Selection

Workload Selection

CPU2026 更系统地使用 multi-workload benchmark。这样做有三个目的:

  1. 覆盖更多 code path 和 application behavior。
  2. 降低针对某个输入的 compiler/hardware optimization 对总分的影响。
  3. 避免 benchmark 被过度“crack”,让优化更偏向通用改进。

代价是 aggregate score 会混合多个行为。

例如 729.abc727.cppcheck 内部同时包含 core-bound/high-IPC 和 memory-bound/low-IPC 子 workload,不能简单地把整个 benchmark 归类为某一种瓶颈。论文用 Basic Block Vector (BBV) recurrence plot 辅助观察不同输入是否触发不同 execution phase,同时在文档中保留自定义 input set 的方法,便于研究者进一步拆解。

Culling Criteria

候选被排除的核心原因包括:

  1. Determinism:不能保证不同平台执行相同 work 的候选被排除。heuristic search、linear solver、gradient descent 等程序容易因微小数值差异走不同路径。
  2. Development divergence:为了 portability 去除 ISA-specific 优化后,如果程序行为不再代表真实部署,则不应进入 suite。
  3. Domain redundancy and scope:同一应用域不能过度集中;用户群太窄、行业代表性不足的候选也会被排除。
  4. Codebase health:偏好现代、活跃维护、结构清晰的代码;过老、混乱、不可维护的代码不适合作为长期 benchmark。
  5. Insufficient maturity:如果候选还不能及时完成跨系统移植和验证,会被推迟到未来 suite。
  6. Excessive I/O or peaky profiles:SPEC CPU 主要测 CPU 和 memory subsystem;大量 I/O 或少数函数支配 runtime 的候选容易引入存储系统影响或窄优化机会。
  7. Potential bias:疑似为某个委员会成员架构预调优的候选会被严格审查。

典型排除案例

论文用若干案例说明选择原则。

Modern AI workload:llama.cppwhisper.cpp 进入了较深评估阶段,但去掉 intrinsic 后变成不真实的 compute-bound hot loop,MPKI 和真实部署差异明显。llama.cpp 还需要用离线 token 序列“固定”推理路径才能验证结果,这进一步偏离真实 inference。

Cryptography:AES、RSA 等生产级 crypto 主要依赖 hand-tuned assembly 和 ISA intrinsic。移除后得到的 generic crypto workload 不代表真实部署。但 750.sealcrypto 被保留,因为 homomorphic encryption 的有限域数学核心仍有通用 CPU 价值。

Media codecs:AV1/AOM、Opus 等候选依赖架构专用实现;Opus 还因处理过快导致 multi-copy SPECrate 变成 I/O-bound。相对地,817.flac 被保留,因为其多线程程序向单个文件写入,不会产生相同 I/O 问题。

Optimization/search workload:737.gmsh 原本有 adaptive mesh refinement,迭代次数会因 FP 数值、compiler flag、ISA 差异而变化;通过禁用 adaptive phase 保留了代表性并获得确定性。HiGHS linear programming solver 则无法在代表性输入上保证 equal work,因此被排除。846.minizinc 使用 unsatisfiable inputs,使 solver 必须穷尽搜索,从而自然保证 equal work。

Compression:xz、brotli、7-zip、zstd 都是高质量候选,但行为相似。最终 777.zstd 进入 intrate,因为应用广泛;801.xz 进入 intspeed,因为它是压缩候选中提供多线程实现的程序。

Longevity and Portability:长期可用性

SPEC CPU suite 会被冻结并使用多年,因此代码必须能跨未来编译器、新系统和旧硬件长期运行。CPU2026 采用 C18、C++17、Fortran 2018,并进行大规模代码硬化。

主要工程工作包括:

  1. Code hardening and standards compliance:清除 undefined behavior、unspecified behavior、implementation-defined behavior;修复危险 pointer conversion、data overflow、未初始化变量、initialization-order bug。
  2. Data model portability:修复整数类型、signed/unsigned、object size、index type 等问题,避免在不同 ABI 上行为变化。
  3. C++17 modernization:替换 std::bind2ndstd::random_shuffleregister 等过时或不合标准的用法,改用现代标准接口。
  4. Removal of non-standard extensions:移除 compiler-specific attribute 和非标准关键字,降低特定编译器偏向。
  5. Endianness:在 IBM AIX/POWER big-endian 系统上验证,修复输入文件格式、byte order、type punning、union bit-field、little-endian pointer assumption 等问题。
  6. Operating systems:支持 Windows x86-64 和 aarch64;利用 MinGW 区分 Windows OS 问题、Windows compiler 问题和 GNU/POSIX dependency。
  7. IO reduction:用 stracesaremon 等分析 I/O;裁剪输入输出、移除周期性写、使用 buffered output、替换 std::endl、合并 open/close。
  8. Memory safety:使用 GCC/LLVM ASan、TSan,以及 AmpereOne 上的 ARM MTE 发现 buffer overrun、use-after-free、data race 等问题。

这些工作不是只服务 SPEC。论文列出多个 upstream patch,例如 nestgem5diamondgmshns3vpr 等项目都从 SPEC CPU development 中获得修复。

Reference System

CPU2026 使用 Lenovo ThinkSystem HR330A 作为参考系统,处理器为 2018 年的 Ampere eMAG 8180 aarch64。

reference machine 的作用是提供 reference time 和 energy,用于 SPEC ratio 归一化。选择较旧机器可以让现代系统 score 通常大于 1;但同一 suite 内两台系统的相对性能不依赖具体 reference machine

Characterization

展示如何用 PMC、Top-down Microarchitectural Analysis、BBV 和 performance plot 理解候选行为

实验平台:

Component Configuration
Processor AMD EPYC 9755, Zen 5
Frequency 2.7 GHz, max boost 4.1 GHz
Memory 2.3 TiB DDR5-6400
Cache L1 32 KiB I + 48 KiB D, L2 1 MiB, L3 512 MiB
Compiler GCC 15.2, -O3
OS Ubuntu 24.04 LTS, Linux 6.8.0-44-generic

SPECrate characterization 使用 single copy,统计 IPC、frontend-bound、backend-bound、bad speculation/lost cycles、retiring fraction。主要观察:

  1. intrate 行为比 fprate 更均衡,既有 frontend-bound,也有 backend-bound。
  2. 727.cppcheck753.ns3709.cactus 等表现出明显 frontend pressure。
  3. 708.sqlite749.fotonik3d765.roms 等更偏 backend-bound。
  4. 777.zstd731.astcenc lost cycles 较高,符合压缩类应用控制流不规则的特点。
  5. fprate 整体更偏 backend-bound,lost cycles 通常更低。

SPECspeed characterization 使用 128 threads,但论文提醒并非所有 speed workload 都能扩展到 128 threads。观察包括:

  1. intspeed 同时包含 frontend-bound 和 backend-bound 行为。
  2. fpspeed 更一致地 backend-limited,反映许多 FP/HPC workload 的内存和执行资源压力。
  3. 800.pot3d801.xz 初步显示较强 contention,值得研究 lock contention、dirty-line sharing、cache coherence 和 snoop filter 行为。

Int Ref Rate:

Fp Ref Rate:

Int Ref Speed:

Fp Ref Speed:

BBV recurrence plot 用于观察 execution phase。论文以 853.ns3 为例展示七个 workload 对应七个主要 block,其中第三个 workload DTLB miss 较高,在 perf plot 中对应 backend bottleneck 上升和 IPC 下降。HPC 类 benchmark 通常自相似性更高、phase 较少;许多非 HPC benchmark 有多个 phase。

RRR:面向异构多程序负载的运行方式

论文提出 Rolling Round-Robin Rate (RRR)。

传统 SPECrate 是 homogeneous capacity benchmark:同一时间运行同一个 component benchmark 的多个 copy。这种方法从 1992 年以来很稳定,适合测系统吞吐,但现代服务器和多租户环境往往同时运行不同 workload。

RRR 的目标是标准化 multiprogrammed workload construction。对于包含 N 个 benchmark 的 suite 和 M 个 core:

  1. 每个 core 按固定顺序执行全部 N 个 benchmark。
  2. 不同 core 的起始 benchmark 以 round-robin 方式错开。例如 Core 0 从 A 开始,Core 1 从 B 开始。
  3. 每个 core 继续执行其 schedule,直到所有 benchmark 在所有 core 上完成。
  4. 每个 benchmark 在每个 core 上都运行一次,从 user-level CPU execution 角度保证相同 instruction count,降低 sample imbalance。
  5. harness 提供 inc 参数和 subset 选择,允许构造不同异构场景。

RRR 目前是 exhibition methodology,因为 scoring 尚未定型。论文没有强行规定 cumulative IPC、average throughput、harmonic mean、fairness index 等指标,而是先提供确定性 workload schedule 和原始执行数据。这样后续研究可以在同一 workload construction 上比较 OS scheduling、resource partitioning、heterogeneous multicore policy 和 multiprogram throughput metric。

RRR 的价值在于把 multiprogramming 研究中常见的 ad-hoc workload mix 和 ad-hoc scheduling policy 标准化。它保留真实复杂应用之间的 microarchitectural interaction,而不是用手工隔离的 kernel 或自定义混合替代完整程序