PhyAI: Unified Inference Engine for Physical AI at Edge and Cloud
Researchers have introduced a new inference engine named PhyAI, designed to enhance the entire process of physical AI policy inference. This includes everything from model assessment to cloud reinforcement learning and GPU serving. PhyAI provides a cohesive runtime that merges graph execution, memory management, and parallel services, addressing the challenge of disjointed inference programs while preserving architecture-specific logic through model adapters. This enables a unified codebase to run both vision-language-action (VLA) models and world-action models (WAMs) on a variety of GPUs, whether onboard, at the edge, or in the cloud. On its launch day, it successfully integrated MiniCPM-Robot through its adapter interface, showing speed boosts of 1.40x-4.65x over existing versions. You can find the detailed paper on arXiv using the identifier 2608.03682.
Key facts
- PhyAI is a Physical AI inference engine unifying inference across model evaluation, cloud RL rollout, edge GPU serving, and onboard deployment.
- It uses a single runtime with model adapters for architecture-specific logic, sharing graph execution, kernels, memory management, and parallel services.
- The same codebase runs VLA models and WAMs on single or multiple GPUs across onboard, edge, and cloud deployments.
- The adapter interface allowed adding MiniCPM-Robot on the day of its release.
- PhyAI achieves 1.40x-4.65x speedups over official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot.
- The paper is available on arXiv with identifier 2608.03682.
- The system is designed for real-time physical AI at the edge and scalable rollouts in the cloud.
Entities
Institutions
- arXiv