PhAIL – Real-robot benchmark for AI models
PhAIL – Real-robot benchmark for AI models
I built this because I couldn't find honest numbers on how well VLA models [1] actually work on commercial tasks. I come from search ranking at Google where you measure everything, and in robotics nobody seemed to know. PhAIL runs four models (OpenPI/pi0.5, GR00T, ACT, SmolVLA) on bin-to-bin order picking – one of the most common warehouse operations. Same robot (Franka FR3), same objects, hundreds of blind runs. The operator doesn't know which model is running. Best model: 64 UPH. Human teleoperating the same robot: 330. Human by hand: 1,300+. Everything is public – every run with synced video and telemetry, the fine-tuning dataset, training scripts. The leaderboard is open for submissions. Happy to answer questions about methodology, the models, or what we observed. [1] Vision-Language-Action: https://en.wikipedia.org/wiki/Vision-language-action_model
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
Replicover – Find the hottest AI models on Replicate
AI Models PAI
Tinx.ai – Prebuilt AI Models for You
Robot and Puppy Codepen
INDI Robot Combat Competition (Dutch)
Teleport: A Telepresence Robot
Homemade Wheeled Biped Robot
Bard – An Experiment in Robot Poetry
My 2 wheeled robot
Robot Rescue Headquarters