RooflineBench: Benchmarking On-Device LLMs with Roofline Analysis
A new benchmarking framework, RooflineBench, has been introduced to evaluate the performance of on-device large language models (LLMs) and small language models (SLMs) on edge hardware. The framework, detailed in a paper on arXiv (2602.11506), uses the Roofline model to unify architectural primitives and hardware constraints through operational intensity (OI). It defines an inference-potential region and introduces a novel metric, Relative Inference Potential, to compare efficiency differences between LLMs on the same hardware. Extensive empirical analysis across diverse compute tiers reveals that performance and OI variations are significantly influenced by sequence length, and a critical regression in OI is identified. The work addresses the challenge of objectively measuring theoretical performance ceilings of diverse architectures on heterogeneous platforms, which is crucial for the transition toward localized intelligence.
Key facts
- Framework named RooflineBench
- Based on Roofline model
- Introduces Relative Inference Potential metric
- Uses operational intensity (OI) as a lens
- Empirical analysis across diverse compute tiers
- Sequence length significantly influences performance and OI
- Identifies critical regression in OI
- Paper on arXiv with ID 2602.11506
Entities
Institutions
- arXiv