# Indus Valley Script

> Built a pipeline that extracts every seal photograph from the Corpus of Indus Seals and Inscriptions and benchmarks VLM transcription against Mahadevan's 417-sign list. Designed the three review tools that let one person verify, correct and label its output at keyboard speed, and feed every judgement back into the pipeline.

- Company: Personal project
- Timeline: Jul 2026 – Aug 2026
- Role: Solo, product and design: wrote the PRD, designed the QC, transcription review and sign labelling tools, built the pipeline with Claude Code, reviewed every crop, ran the benchmark
- Team: Solo
- Key outcome: 3,826 human-verified seal crops with a reproducible public manifest, reviewed at one keystroke per crop. Four rounds of review feedback rewrote the extractor. The benchmark showed that sign identification, not scan resolution or reading order, is the bottleneck, and re-scoped the pipeline to segment, match, then disambiguate.

---

![Pipeline from the CISI scan to verified transcriptions](/case-studies/11-indus-script/01-pipeline.svg)

*Code, coordinates and transcription facts are published. The photographs stay on local disk.*

## The Problem

The Indus script has no open, machine-readable corpus.

- Mahadevan's 1977 concordance covers ~2,900 texts, as a printed book
- ICIT covers ~4,660 artefacts, behind a login with no export
- The only open dataset has 179 artefacts, about 5% of known inscriptions
- No labelled image dataset of the seals exists anywhere

The source photographs are copyrighted. So the project ships code, coordinates and transcription facts, and anyone can fetch the scans from archive.org and rebuild the dataset with one command.

## Design Principle

Every stage of the pipeline produces output a machine cannot fully trust. So every stage got a tool where a human judgement costs one keystroke, and every judgement is written back into the data the pipeline runs on. The reviewer never fills a form. They look, press a key, and move on.

## Extraction

CISI Volume 1 has ~3,600 seal photographs on 364 plate pages, scanned at 100 dpi. Every photograph has a caption printed beneath it.

- Apple Vision OCR reads the captions; a strict grammar parses them into labels
- Each photograph attaches to the caption below it and is cropped at native resolution
- The manifest records label, page, bounding box and checksum

| | |
|---|---|
| Crops | 3,826 |
| Auto-labelled correctly | 97.8% |
| Reviewed by hand | All, over four rounds |
| Median crop | 208 × 164 px |

## QC Review Tool

![QC review tool showing the page scan beside the crop for M-8A](/case-studies/11-indus-script/02-qc-review.webp)

*Page scan left with the crop outlined, the crop and its label right. The verdict buttons carry their keys.*

- One verdict per key: space for correct, L for a wrong label, C for a wrong crop, B for both, U for unsure
- E opens crop editing: drag the handles or nudge an edge with the arrow keys, then Enter re-cuts the crop from the PDF and rewrites the manifest
- Filters for undecided, flagged, marked bad and a review list, plus jump to any label
- Manifest checks (OCR failure, duplicate label, out of sequence, size outlier) show as badges on the item
- A progress bar across the top. Every decision auto-saves; nothing is submitted

The verdicts fed the extractor directly.

| Round | Reviewer input | Pipeline change |
|---|---|---|
| 1 | ~160 verdicts on a test run | Extractor rewritten; matched 47 of 47 label corrections |
| 2 | 37 crops marked wrong | Three pixel checks added; ~190 crops re-cut across the volume |
| 3 | Seals reported with no crop | Caption parser fixed for OCR homoglyphs; 5 hidden seals recovered |
| 4 | 11 crops marked too tight | Edge-growth rule replaced the mean-intensity test |

## Signary

![Signary reference sheet showing Mahadevan signs 1 to 90](/case-studies/11-indus-script/03-signary-sheet.webp)

*Sheet 1 of 5. Every sign carries its Mahadevan 1977 number; sign 342, the jar, anchors the numbering.*

## Benchmark

Pack 1: 20 Mohenjo-daro unicorn seals with a single clear sign row, the easy case. A VLM read each crop against the five signary sheets and returned a sequence with a confidence per sign.

![Transcription review page for pack 1](/case-studies/11-indus-script/04-transcription-review.webp)

*Signs are drawn as glyphs, not numbers, laid out right to left to mirror the seal. Confidence is colour coded. Sort by least confident first, or show only seals below 0.35.*

| | |
|---|---|
| Sign positions | 121 |
| Illegible | 7 |
| Mean confidence | 0.34 |
| Jar sign share | 2.5%, expected ~10% |

The jar sign check failed. Two independent reads of the same seal shared no signs at all. The model could see the strokes but could not resolve them to one of 417 numbered cells, because whole families of signs differ by one stroke or one tick.

## Ground Truth

![Seal M-1a in the labelling tool](/case-studies/11-indus-script/05-label-tool-m1a.webp)

*Seal M-1a at native resolution. The strokes are countable, so scan resolution was not the bottleneck.*

Hand transcribing M-8a gave 244 17 336. The VLM had read 246 151 340: correct order, correct segmentation, wrong IDs, and two of the three were adjacent members of the right family. A local chamfer matcher against the 417 templates put the true sign at rank 1, rank 1 and rank 3.

| Position | Truth | VLM | Template match |
|---|---|---|---|
| 1 | 244 | 246 | 244, rank 1 |
| 2 | 17 | 151 | 17, rank 1 |
| 3 | 336 | 340 | 336, rank 3 |

## Labelling Tool

The project needed two datasets: sign IDs to score the transcriber, and sign boxes to train the segmentation step. The tool was designed so one gesture yields both.

![Drawing a box around a sign on M-8a](/case-studies/11-indus-script/06-label-tool-box.webp)

*Drag a box around a sign. The matcher runs on the box in 3 ms and returns the eight nearest templates as glyphs; keys 1 to 8 pick one.*

- Box first, identify second. The box is the segmentation label; the ID is the classification label
- Candidates are shown as glyphs so the reviewer compares shapes, not numbers. A typed M77 number is the fallback
- Boxes number themselves right to left and the sequence panel mirrors the photograph, so nothing is reversed in the reviewer's head
- The earlier VLM reading stays in view, so each label is also a scored comparison
- A notes field for damage or a second line; a done checkbox and a count for the pack
- Every edit auto-saves. Tab cycles boxes, Del removes one, arrows change seal

![Candidate grid for the first sign of M-1a](/case-studies/11-indus-script/07-label-tool-candidates.webp)

*Sign 397 at rank 1 for the first sign of M-1a, agreeing with the VLM's most confident read.*

![Three labelled boxes on M-8a](/case-studies/11-indus-script/08-label-tool-sequence.webp)

*M-8a fully labelled: 244, 17, 336. The sequence panel reads as the photograph does.*

## What Changed

The benchmark re-scoped the pipeline.

1. Segment the inscription band into sign boxes. This is the open problem: the signs abut, and projection heuristics fail. The labelling tool's boxes are the training data
2. Classify each box locally against the 417 templates and return a shortlist
3. Ask the VLM to pick from the shortlist, a 5-way choice instead of a 417-way one

---

Other pages: [Home](https://rupakmishra.com/index.md) · [Work](https://rupakmishra.com/work.md) · [About](https://rupakmishra.com/about.md) · [Résumé](https://rupakmishra.com/resume.md)

Source: https://rupakmishra.com/work/indus-script
