Pipeline
Audio → VAD/endpointing → streaming Urdu STT → normalization/entities
→ reasoning workflow → approved Urdu text → streaming TTS
A cascaded pipeline preserves a text trail and allows separate speech benchmarking and provider replacement.
Candidates
Benchmark Deepgram Nova-3 Urdu, Azure Speech Pakistani Urdu and one self-hosted candidate on real Punjab speech.
STT metrics
WER, entity error, chemical/crop/variety accuracy, numbers/units, code-switching, interim/final latency, background noise, accent/gender and low-cost microphone slices.
Dynamic terminology
Pass crop, district/villages, varieties, likely conditions, trade names, active ingredients and Urdu/Roman-Urdu/English spellings where supported. Confirm low-confidence chemical, quantity and unit entities.
Normalization
Retain original transcript, normalized Urdu, Roman Urdu/English spans, entities, N-best alternatives, resolved units and confirmation-required flags.
Turn handling
Compare tuned VAD, STT endpointing, push-to-talk and transcript-aware turn detection only after Urdu field evaluation.
TTS
Benchmark Pakistani Urdu naturalness, intelligibility, numbers, active ingredients, trade names and first-byte latency. Maintain pronunciation dictionaries and provider-specific SSML adapters.
Voice inside the live call
Audio remains continuous while the call is connected. Streaming STT produces interim text for UI only and finalized turns for reasoning. TTS supports barge-in, and camera guidance is synchronized with concise spoken Urdu. Network degradation prioritizes two-way audio before video.