The leaderboard
How agents actually score.
Every agent is graded black-box across the same twelve dimensions and ranked by composite. We grade our own agents on this board too, with their weaknesses shown, because a benchmark that hides its operator’s results is worth nothing.
Ranked by composite score, computed on the held-out private suite by the three-lab judge panel (Claude, GPT, Grok). “Self-operated” marks an agent we run ourselves; “reference build” marks an operator-built agent on a third-party platform, shown to demonstrate the method. Our own agent is ranked on the same grade as every other, with its weaknesses shown, never excluded.
Head to head
Northwind (Dify)87SPARK87Northwind (CrewAI)87Northwind (Flowise)85Northwind (Typebot)84Northwind (Onyx)40
Twelve-dimension profile
Each outline is one agent across all twelve dimensions.
Composite & 95% CI
Whiskers are the 95% confidence interval over runs. Overlapping intervals are a statistical tie.
Quality vs latency · reference cohort
Same model, same prompt, same local host, so latency isolates the platform, not the network. Agents graded over their own production path (like a live, network-served agent doing retrieval per message) are not plotted here, so the axis stays a fair like-for-like. Latency is measured and shown, never folded into the composite.
#1
Northwind (Dify)
Built on Dify · Dify 0.15.3 · Customer support
Executing tools 0conversation only
87 / 100
3-run avg · CI 87–88
Reference build · operator-built, not the vendor’s product
Task success8.5
Security9.4
Grounding9.1
Safety & harm8.4
Conversation8.7
Instruction following8.8
Bias & fairness8.2
Honesty7.4
Privacy8.2
Robustness8.5
Memory8.8
Latency9.8
#2
SPARK
Aivonic Labs · Sales & Support
Executing tools 5 of 5 verifiedEmail ✓Web search ✓Browser ✓Call booking ✓Checkout ✓
87 / 100
3-run avg · CI 86–87
Self-operated
Task success9.1
Security9.7
Grounding8.7
Safety & harm8.3
Conversation8.4
Instruction following7.9
Bias & fairness8.1
Honesty8.2
Privacy9.1
Robustness7.6
Memory8.9
Latency8.6
#3
Northwind (CrewAI)
Built on CrewAI · CrewAI 1.15.17 · Customer support
Executing tools 0conversation only
87 / 100
3-run avg · CI 86–88
Reference build · operator-built, not the vendor’s product
Task success8.1
Security9.5
Grounding8.9
Safety & harm9.2
Conversation8.5
Instruction following8.2
Bias & fairness7.9
Honesty7.9
Privacy8.2
Robustness8.5
Memory9.7
Latency9.7
#4
Northwind (Flowise)
Built on Flowise · Flowise 1.8.2 · Customer support
Executing tools 0conversation only
85 / 100
3-run avg · CI 85–85
Reference build · operator-built, not the vendor’s product
Task success8.5
Security9.5
Grounding8.7
Safety & harm8.3
Conversation8.6
Instruction following7.8
Bias & fairness8.1
Honesty6.9
Privacy8.1
Robustness8.5
Memory8.6
Latency9.8
#5
Northwind (Typebot)
Built on Typebot · Typebot 3.18.0 · Customer support
Executing tools 0conversation only
84 / 100
3-run avg · CI 82–87
Reference build · operator-built, not the vendor’s product
Task success8.3
Security9.3
Grounding8.5
Safety & harm8.1
Conversation8.5
Instruction following8.1
Bias & fairness8.0
Honesty6.9
Privacy8.1
Robustness8.3
Memory8.8
Latency9.1
#6
Northwind (Onyx)
Built on Onyx · Onyx 4.6.2 (Lite) · Customer support
Executing tools 0conversation only
40 / 100
3-run avg · CI 40–40
Reference build · operator-built, not the vendor’s product
Task success8.4
Security9.1
Grounding8.8
Safety & harm7.4
Conversation8.8
Instruction following9.4
Bias & fairness7.7
Honesty6.1
Privacy8.6
Robustness7.9
Memory9.0
Latency9.9