Rippling Tests 15 AI Models on Real Payroll Data: Cheapest Ties Most Expensive
The Gist
- Rippling ran 2,100 scored agent runs per model on real payroll data
- Opus 4.6 tied GPT-5.5 med with a 91% pass rate for $1,453
- GPT-5.5 low outperformed on speed, costing $1,435 with 89.5% pass rate
- Slowest 10% task completion times were key customer experience metrics
Key Quotes
The cheaper models actually did more back-and-forth, not less. They aren’t cutting corners. They’re just priced differently.
Checking the work is the product. No model on this list saves you, including the $4,359 one.
Key Insights
- Cheaper AI models often match or exceed the performance of more expensive ones, with GLM 5.2 saving $687 compared to GPT-5.5 low at identical accuracy.
- The accuracy spread across the leading AI models is narrow, with seven models landing between 88.5% and 89.5%, and the leader at 91.0%.
- Tuning AI models and instructions can significantly impact accuracy, with Rippling's tuning efforts estimated to be worth 1-2 points of accuracy.
- Newer AI models are not always better, as Grok 4.6 was worse and slower than Grok 4.5 in Rippling's tests.
- AI model performance should be evaluated based on specific tasks, as cheaper models may excel in non-time-sensitive tasks while more expensive models are better for live customer interactions.
- Verification of AI outputs is critical, as models can report incorrect answers as correct, posing significant risks in production environments.
Actionable Takeaways
- Evaluate AI models based on specific use cases and re-run tests before upgrading to newer versions.
- Invest in tuning AI models and instructions to maximize accuracy and performance.
- Monitor AI costs closely, as cheaper models can provide significant margin improvements without sacrificing accuracy.
- Implement robust verification processes to ensure AI outputs are correct, especially for critical tasks.
Data Points
- 91.0% (The highest accuracy achieved by the best-tuned AI model in Rippling's study.)
- 88.5% to 89.5% (The accuracy range for seven AI models tested by Rippling.)
- $621 (The cost of GLM 5.2, which achieved 88.7% accuracy without tuning.)
- $2,509 (The cost of Opus 5, which performed similarly to Grok 4.5 but was significantly more expensive.)
- 71 seconds to 131 seconds (The increase in response time when comparing Grok 4.5 to Grok 4.6.)
RevBots.ai View:
AI model performance benchmarks in production systems reveal cost and accuracy trade-offs that matter for GTM efficiency.
Full Story:
SaaStr →
Join The RevBots ARMy
The insider daily for Autonomous Revenue Masters.