Evaluation principle
Evaluate the complete live decision process, not only model accuracy. Thresholds are approved before farmer exposure and stratified by crop, condition, district, accent, device, network and risk.
Test layers
- Unit and schema tests.
- WebRTC room/token/webhook contract tests.
- Network simulation: packet loss, jitter, bandwidth changes, reconnect and TURN.
- Speech, frame-quality and camera-guidance tests.
- Retrieval/research and reasoning cases.
- Safety, prompt-injection and policy tests.
- End-to-end live Urdu simulations.
- Shadow evaluation and controlled field pilot.
Urdu speech set
Cover genders, ages, Punjab regions, noise, machinery, wind, echo, low-cost microphones, code switching, village/crop/chemical names, quantities and units. Measure word error, critical-entity error, latency and confirmation burden.
Live media and camera-guidance set
Include expert-confirmed cases, healthy controls, lookalikes, stages/varieties, unseen farms, target phone types, lighting, motion blur, low bandwidth, camera rotation, background clutter, invalid media and spray/environmental injury.
- Call setup and reconnect success.
- Useful-frame acquisition and rejected-frame precision.
- Number and clarity of camera instructions.
- Time to evidence sufficiency.
- Failure to notice loss of usable visual evidence.
- Selected-frame diagnosis performance versus full reviewed video.
Retrieval and research set
Test exact terms, multilingual concepts, hierarchy, current versus expired material, local applicability, conflicting research, citation resolution, authority classification and “no sufficient source” behaviour.
Reasoning metrics
- Primary/differential hypotheses and evidence coverage.
- Value of selected next question or camera objective.
- Contradiction detection and local applicability.
- Unnecessary research/capture and premature stopping.
- Calibration, citation correctness and expert explanation quality.
Safety metrics
- Unsupported treatment/dose and missed verification.
- Missed escalation or officer-join trigger.
- False certainty and harmful omission.
- Expired/withdrawn rule execution.
- Prompt-injection and adversarial media success.
- Unintended recording or unauthorized room participation.
User and operational metrics
Farmer comprehension, live-call completion, abandonment, network degradation, camera guidance success, time to useful response, officer agreement/join time, queue SLA and cost per completed advisory.
Threshold governance
Set separate go/no-go values for low- and high-risk functions. A high average score cannot compensate for failure on a critical subgroup, unsupported device/network tier or privacy control.