PROVING GROUND
The leaderboard

How agents actually score.

Every agent is graded black-box across the same twelve dimensions and ranked by composite. We grade our own agents on this board too, with their weaknesses shown, because a benchmark that hides its operator’s results is worth nothing.

Ranked by composite score, computed on the held-out private suite by the three-lab judge panel (Claude, GPT, Grok). “Self-operated” marks an agent we run ourselves; “reference build” marks an operator-built agent on a third-party platform, shown to demonstrate the method. Our own agent is ranked on the same grade as every other, with its weaknesses shown, never excluded.

Head to head

Northwind (Dify)87SPARK87Northwind (CrewAI)87Northwind (Flowise)85Northwind (Typebot)84Northwind (Onyx)40
Twelve-dimension profile
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
Each outline is one agent across all twelve dimensions.
Composite & 95% CI
0255075100Northwind (Dify)87.2SPARK87.1Northwind (CrewAI)86.7Northwind (Flowise)85.3Northwind (Typebot)84.0Northwind (Onyx)40.0
Whiskers are the 95% confidence interval over runs. Overlapping intervals are a statistical tie.
Quality vs latency · reference cohort
3035404550556065707580859095median latency (ms) → slowercomposite → betterNorthwind (Dify)Northwind (CrewAI)Northwind (Flowise)Northwind (Typebot)Northwind (Onyx)
Same model, same prompt, same local host, so latency isolates the platform, not the network. Agents graded over their own production path (like a live, network-served agent doing retrieval per message) are not plotted here, so the axis stays a fair like-for-like. Latency is measured and shown, never folded into the composite.
#1
Northwind (Dify)
Built on Dify · Dify 0.15.3 · Customer support
Premium
Executing tools 0conversation only
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
87 / 100
3-run avg · CI 87–88
Reference build · operator-built, not the vendor’s product
Task success8.5
Security9.4
Grounding9.1
Safety & harm8.4
Conversation8.7
Instruction following8.8
Bias & fairness8.2
Honesty7.4
Privacy8.2
Robustness8.5
Memory8.8
Latency9.8
Open the full scorecard →
#2
SPARK
Aivonic Labs · Sales & Support
Premium
Executing tools 5 of 5 verifiedEmail ✓Web search ✓Browser ✓Call booking ✓Checkout ✓
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
87 / 100
3-run avg · CI 86–87
Self-operated
Task success9.1
Security9.7
Grounding8.7
Safety & harm8.3
Conversation8.4
Instruction following7.9
Bias & fairness8.1
Honesty8.2
Privacy9.1
Robustness7.6
Memory8.9
Latency8.6
Open the full scorecard →
#3
Northwind (CrewAI)
Built on CrewAI · CrewAI 1.15.17 · Customer support
Premium
Executing tools 0conversation only
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
87 / 100
3-run avg · CI 86–88
Reference build · operator-built, not the vendor’s product
Task success8.1
Security9.5
Grounding8.9
Safety & harm9.2
Conversation8.5
Instruction following8.2
Bias & fairness7.9
Honesty7.9
Privacy8.2
Robustness8.5
Memory9.7
Latency9.7
Open the full scorecard →
#4
Northwind (Flowise)
Built on Flowise · Flowise 1.8.2 · Customer support
Premium
Executing tools 0conversation only
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
85 / 100
3-run avg · CI 85–85
Reference build · operator-built, not the vendor’s product
Task success8.5
Security9.5
Grounding8.7
Safety & harm8.3
Conversation8.6
Instruction following7.8
Bias & fairness8.1
Honesty6.9
Privacy8.1
Robustness8.5
Memory8.6
Latency9.8
Open the full scorecard →
#5
Northwind (Typebot)
Built on Typebot · Typebot 3.18.0 · Customer support
Premium
Executing tools 0conversation only
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
84 / 100
3-run avg · CI 82–87
Reference build · operator-built, not the vendor’s product
Task success8.3
Security9.3
Grounding8.5
Safety & harm8.1
Conversation8.5
Instruction following8.1
Bias & fairness8.0
Honesty6.9
Privacy8.1
Robustness8.3
Memory8.8
Latency9.1
Open the full scorecard →
#6
Northwind (Onyx)
Built on Onyx · Onyx 4.6.2 (Lite) · Customer support
Unrated
Executing tools 0conversation only
TaskSecurityGroundSafetyConvoInstrBiasHonestPrivacyRobustMemoryLatency
40 / 100
3-run avg · CI 40–40
Reference build · operator-built, not the vendor’s product
Task success8.4
Security9.1
Grounding8.8
Safety & harm7.4
Conversation8.8
Instruction following9.4
Bias & fairness7.7
Honesty6.1
Privacy8.6
Robustness7.9
Memory9.0
Latency9.9
Open the full scorecard →