KKissan Ki Pehchan
Delivery

Cost, Capacity and Scalability

Cost and capacity model for live media, selected-frame reasoning, research, human review and production growth.

BlueprintVersion 0.25 Aug 2026

Demand inputs

InputWhy it matters
Active farmers and sessionsBase demand
Peak concurrent live rooms and durationSFU/TURN, STT/TTS and worker capacity
Video bitrate and TURN relay shareMedia egress cost
Selected frames/stills per caseVision/model/storage cost
Research trigger rateSearch latency and cost
Long-context usageReasoning cost
Officer join/review minutesHuman capacity
Seasonal peak and retriesHeadroom and hidden cost

Cost per completed advisory

Include media transport/TURN, STT minutes, TTS characters, frame extraction and model input, research, storage, databases, monitoring, support and human-review time. Report separately for full-video, reduced-video and fallback sessions.

Cost controls

  • Adaptive subscription and frame sampling rather than full-frame inference.
  • Stop sampling after evidence is accepted.
  • Ask a high-value question before expensive deep research.
  • Cache public versioned low-risk explanations.
  • Use lower-cost models for frame quality and bounded extraction.
  • Reserve strongest reasoning for materially complex cases.
  • Avoid default full-call recording and long retention.

Scaling triggers

TriggerCandidate response
SFU/TURN saturationAdditional nodes, regional TURN and autoscaling
Agent/frame worker saturationIndependent pools
PostgreSQL retrieval contentionDedicated retrieval service/Qdrant
Background backlogDurable queue/event platform
Complex resumable human workflowsLangGraph/equivalent
Availability commitmentsKubernetes HA and DR environment

Scenarios

Build controlled-pilot, district-scale and province-scale models with explicit concurrency, relay percentage, session duration, selected-frame count, research rate, officer staffing and acceptable cost per completed advisory.