Accuracy
- Measure
- share of eligible runs reaching the confirmed final state
- Decision rule
- Did the system reach the task's declared final state, inside the confirmation boundary, without an unrecoverable error?
The world’s first public benchmark for end-to-end consumer-agent execution across 6 localized markets and everyday consumer domains.
MICA evaluates whether a versioned consumer-agent system can complete consumer tasks end to end across six localized markets.
Preview how results will be publishedHow MICA measuresRead the submission requirements
Counted from the canonical records at build time.
The MICA capability framework
The evaluation target is not an individual model but an assembled consumer-agent system spanning understanding, orchestration, model routing, memory, tool use, localisation, safety and recovery.
Test
Exercise the system under local language, identity, payment, permission, handoff and recovery conditions, with realistic user context and changing constraints.
Measure
Record accuracy, speed and cost separately, and do not treat an execution that misses its declared final state as a partial success.
Explain
Trace outcomes across seven system-capability axes without folding diagnosis into the score.
MICA score = 100 × normalized Accuracy × normalized Speed × normalized Cost
All three axes remain published in their own units. The derived score is not a headline result or ranking key until its references, repeat-run contract and uncertainty are calibrated in the pilot.
10 everyday task domains are the proving ground
100 canonical tasks across 10 task domains and 6 localized markets form 600 planned task-market cells. This is evaluation scope, not a measured result.
Why MICA exists
WebArena, WebVoyager, WebShop, AndroidWorld, AppWorld and τ-bench measure end-state task completion. MICA extends that work to assembled consumer-agent systems completing ordinary errands in localized markets with their own channels, identity checks and payment rails.
6 market editions
What we ask the agent to do
Each family is defined by a declared final state and a confirmation boundary the system must not cross without consent.
What is built, and what is not
MICA now provides the task taxonomy, measurement method, publication gate and evidence model. A result appears only after an independent rerun clears the gate.
No verified system results have been published yet.
10 evaluation families · 0 published result families