Code, coordinates and transcription facts are published. The photographs stay on local disk.
The Problem
The Indus script has no open, machine-readable corpus.
- Mahadevan's 1977 concordance covers ~2,900 texts, as a printed book
- ICIT covers ~4,660 artefacts, behind a login with no export
- The only open dataset has 179 artefacts, about 5% of known inscriptions
- No labelled image dataset of the seals exists anywhere
The source photographs are copyrighted. So the project ships code, coordinates and transcription facts, and anyone can fetch the scans from archive.org and rebuild the dataset with one command.
Design Principle
Every stage of the pipeline produces output a machine cannot fully trust. So every stage got a tool where a human judgement costs one keystroke, and every judgement is written back into the data the pipeline runs on. The reviewer never fills a form. They look, press a key, and move on.
Extraction
CISI Volume 1 has ~3,600 seal photographs on 364 plate pages, scanned at 100 dpi. Every photograph has a caption printed beneath it.
- Apple Vision OCR reads the captions; a strict grammar parses them into labels
- Each photograph attaches to the caption below it and is cropped at native resolution
- The manifest records label, page, bounding box and checksum
| Crops | 3,826 |
| Auto-labelled correctly | 97.8% |
| Reviewed by hand | All, over four rounds |
| Median crop | 208 × 164 px |
QC Review Tool

Page scan left with the crop outlined, the crop and its label right. The verdict buttons carry their keys.
- One verdict per key: space for correct, L for a wrong label, C for a wrong crop, B for both, U for unsure
- E opens crop editing: drag the handles or nudge an edge with the arrow keys, then Enter re-cuts the crop from the PDF and rewrites the manifest
- Filters for undecided, flagged, marked bad and a review list, plus jump to any label
- Manifest checks (OCR failure, duplicate label, out of sequence, size outlier) show as badges on the item
- A progress bar across the top. Every decision auto-saves; nothing is submitted
The verdicts fed the extractor directly.
| Round | Reviewer input | Pipeline change |
|---|---|---|
| 1 | ~160 verdicts on a test run | Extractor rewritten; matched 47 of 47 label corrections |
| 2 | 37 crops marked wrong | Three pixel checks added; ~190 crops re-cut across the volume |
| 3 | Seals reported with no crop | Caption parser fixed for OCR homoglyphs; 5 hidden seals recovered |
| 4 | 11 crops marked too tight | Edge-growth rule replaced the mean-intensity test |
Signary

Sheet 1 of 5. Every sign carries its Mahadevan 1977 number; sign 342, the jar, anchors the numbering.
Benchmark
Pack 1: 20 Mohenjo-daro unicorn seals with a single clear sign row, the easy case. A VLM read each crop against the five signary sheets and returned a sequence with a confidence per sign.

Signs are drawn as glyphs, not numbers, laid out right to left to mirror the seal. Confidence is colour coded. Sort by least confident first, or show only seals below 0.35.
| Sign positions | 121 |
| Illegible | 7 |
| Mean confidence | 0.34 |
| Jar sign share | 2.5%, expected ~10% |
The jar sign check failed. Two independent reads of the same seal shared no signs at all. The model could see the strokes but could not resolve them to one of 417 numbered cells, because whole families of signs differ by one stroke or one tick.
Ground Truth

Seal M-1a at native resolution. The strokes are countable, so scan resolution was not the bottleneck.
Hand transcribing M-8a gave 244 17 336. The VLM had read 246 151 340: correct order, correct segmentation, wrong IDs, and two of the three were adjacent members of the right family. A local chamfer matcher against the 417 templates put the true sign at rank 1, rank 1 and rank 3.
| Position | Truth | VLM | Template match |
|---|---|---|---|
| 1 | 244 | 246 | 244, rank 1 |
| 2 | 17 | 151 | 17, rank 1 |
| 3 | 336 | 340 | 336, rank 3 |
Labelling Tool
The project needed two datasets: sign IDs to score the transcriber, and sign boxes to train the segmentation step. The tool was designed so one gesture yields both.

Drag a box around a sign. The matcher runs on the box in 3 ms and returns the eight nearest templates as glyphs; keys 1 to 8 pick one.
- Box first, identify second. The box is the segmentation label; the ID is the classification label
- Candidates are shown as glyphs so the reviewer compares shapes, not numbers. A typed M77 number is the fallback
- Boxes number themselves right to left and the sequence panel mirrors the photograph, so nothing is reversed in the reviewer's head
- The earlier VLM reading stays in view, so each label is also a scored comparison
- A notes field for damage or a second line; a done checkbox and a count for the pack
- Every edit auto-saves. Tab cycles boxes, Del removes one, arrows change seal

Sign 397 at rank 1 for the first sign of M-1a, agreeing with the VLM's most confident read.

M-8a fully labelled: 244, 17, 336. The sequence panel reads as the photograph does.
What Changed
The benchmark re-scoped the pipeline.
- Segment the inscription band into sign boxes. This is the open problem: the signs abut, and projection heuristics fail. The labelling tool's boxes are the training data
- Classify each box locally against the 417 templates and return a shortlist
- Ask the VLM to pick from the shortlist, a 5-way choice instead of a 417-way one