KKissan Ki Pehchan
AI and Knowledge

Knowledge Ingestion and Lifecycle

Document submission, parsing, hierarchy, metadata, approval, publication and withdrawal.

BlueprintVersion 0.25 Aug 2026

Lifecycle

Submit → malware scan → classify → authority review
→ version/effective dates → parse layout/tables
→ build hierarchy → generate navigation summaries
→ metadata/entities → index → quality check → publish
→ monitor → expire/withdraw

Preserve structure

Document identity, heading tree, sections, paragraphs, tables, figures/captions, footnotes, appendices and source anchors.

Generated chapter summaries are labelled derived_navigation, never original authority.

Chunking

Use coherent semantic units, not only fixed token sizes. Link each chunk to parent section, neighbours, document version, entities, authority and effective/expiry dates.

Metadata

Source ID, title, issuer, version, effective/expiry dates, approval, crop, variety, district/province, season, stage, condition, active ingredient, document type, language, page/section and sensitivity.

Quality

Validate heading tree, tables, Urdu encoding, figure links, sample retrieval queries, authority/dates and summary faithfulness.

Withdrawal

Immediately exclude the version from operational retrieval, invalidate relevant caches and create an audit event. Historical cases retain the version used.