How AI extracts parties, survey numbers, extents and charges from Tamil and English land records, ECs, patta, FMB, and turns them into a verified, structured title chain.
A Tamil Nadu title file is not one document. It is a chain of registered deeds from the Sub-Registrar, encumbrance certificate entries spanning decades, patta and chitta from the Revenue department, an FMB sketch from Survey and Settlement, and planning permissions from DTCP or CMDA. The formats differ, the languages mix Tamil and English, and the copies are often scanned images of handwritten registers.
A diligent human reader can process this, but slowly and unevenly. The practical consequence for a corporate buyer is a diligence window measured in weeks, and a quality level that depends on which junior read which page on which day.
The model reads each scanned page and pulls out the fields that matter: the parties to each transaction, the date and document number, the survey number and extent, the consideration, and the nature of the entry, sale, mortgage, release, partition, gift or attachment. Tamil-language deeds and registers are read natively rather than skipped or sent for separate translation.
The output is a structured title chain: every transaction on the parcel in order, every charge with its creation and release, every extent stated by every record. That structure is what makes verification possible at scale, because software can now compare fields across records instead of a person flipping between PDFs.
With the chain structured, the 30-point verification runs as queries: does the seller's name match the last recorded transferee, does the patta extent agree with the FMB and the deed, is every mortgage matched by a registered release, does any entry indicate a pending attachment or lis pendens. Each failed check becomes a finding graded Clear to Critical, citing the exact entry behind it.
This is the part manual diligence most often gets wrong, not because lawyers cannot do it, but because exhaustively cross-referencing fifty years of entries against four other record types is exactly the kind of work humans skip under deadline pressure.
Extraction accuracy is high but not perfect, and source registers themselves contain errors. Anything that affects a decision is verified by a person against the original image, and the legal conclusions, marketability, the effect of an old partition, the strength of a possession claim, are validated by independent counsel before a report is relied on.
The honest framing is leverage, not replacement: the machine reads everything so the expert can judge the few entries that matter.
Yes. Modern models read Tamil script in scanned deeds and revenue registers alongside English, extracting parties, dates, survey numbers and extents. Accuracy is verified by a human reviewer against the original image for any entry that affects a finding.
Registered deeds, encumbrance certificate entries, patta, chitta, A-Register extracts and FMB sketches. Together they yield a structured chain of every transaction, charge and extent on the parcel over 50 or more years.
No. It feeds it. The structured chain and graded findings are validated by independent counsel, who issues the marketability opinion. Extraction removes the transcription work, not the judgement.