CASE STUDY

Indus Valley Script

Built a pipeline that extracts every seal photograph from the Corpus of Indus Seals and Inscriptions and benchmarks VLM transcription against Mahadevan's 417-sign list. Designed the three review tools that let one person verify, correct and label its output at keyboard speed, and feed every judgement back into the pipeline.

Company
Personal project
Timeline
Jul 2026 – Aug 2026
My Role
Solo, product and design: wrote the PRD, designed the QC, transcription review and sign labelling tools, built the pipeline with Claude Code, reviewed every crop, ran the benchmark
Team
Solo
Key Outcome
3,826 human-verified seal crops with a reproducible public manifest, reviewed at one keystroke per crop. Four rounds of review feedback rewrote the extractor. The benchmark showed that sign identification, not scan resolution or reading order, is the bottleneck, and re-scoped the pipeline to segment, match, then disambiguate.

Pipeline from the CISI scan to verified transcriptions

Code, coordinates and transcription facts are published. The photographs stay on local disk.

The Problem

The Indus script has no open, machine-readable corpus.

The source photographs are copyrighted. So the project ships code, coordinates and transcription facts, and anyone can fetch the scans from archive.org and rebuild the dataset with one command.

Design Principle

Every stage of the pipeline produces output a machine cannot fully trust. So every stage got a tool where a human judgement costs one keystroke, and every judgement is written back into the data the pipeline runs on. The reviewer never fills a form. They look, press a key, and move on.

Extraction

CISI Volume 1 has ~3,600 seal photographs on 364 plate pages, scanned at 100 dpi. Every photograph has a caption printed beneath it.

Crops3,826
Auto-labelled correctly97.8%
Reviewed by handAll, over four rounds
Median crop208 × 164 px

QC Review Tool

QC review tool showing the page scan beside the crop for M-8A

Page scan left with the crop outlined, the crop and its label right. The verdict buttons carry their keys.

The verdicts fed the extractor directly.

RoundReviewer inputPipeline change
1~160 verdicts on a test runExtractor rewritten; matched 47 of 47 label corrections
237 crops marked wrongThree pixel checks added; ~190 crops re-cut across the volume
3Seals reported with no cropCaption parser fixed for OCR homoglyphs; 5 hidden seals recovered
411 crops marked too tightEdge-growth rule replaced the mean-intensity test

Signary

Signary reference sheet showing Mahadevan signs 1 to 90

Sheet 1 of 5. Every sign carries its Mahadevan 1977 number; sign 342, the jar, anchors the numbering.

Benchmark

Pack 1: 20 Mohenjo-daro unicorn seals with a single clear sign row, the easy case. A VLM read each crop against the five signary sheets and returned a sequence with a confidence per sign.

Transcription review page for pack 1

Signs are drawn as glyphs, not numbers, laid out right to left to mirror the seal. Confidence is colour coded. Sort by least confident first, or show only seals below 0.35.

Sign positions121
Illegible7
Mean confidence0.34
Jar sign share2.5%, expected ~10%

The jar sign check failed. Two independent reads of the same seal shared no signs at all. The model could see the strokes but could not resolve them to one of 417 numbered cells, because whole families of signs differ by one stroke or one tick.

Ground Truth

Seal M-1a in the labelling tool

Seal M-1a at native resolution. The strokes are countable, so scan resolution was not the bottleneck.

Hand transcribing M-8a gave 244 17 336. The VLM had read 246 151 340: correct order, correct segmentation, wrong IDs, and two of the three were adjacent members of the right family. A local chamfer matcher against the 417 templates put the true sign at rank 1, rank 1 and rank 3.

PositionTruthVLMTemplate match
1244246244, rank 1
21715117, rank 1
3336340336, rank 3

Labelling Tool

The project needed two datasets: sign IDs to score the transcriber, and sign boxes to train the segmentation step. The tool was designed so one gesture yields both.

Drawing a box around a sign on M-8a

Drag a box around a sign. The matcher runs on the box in 3 ms and returns the eight nearest templates as glyphs; keys 1 to 8 pick one.

Candidate grid for the first sign of M-1a

Sign 397 at rank 1 for the first sign of M-1a, agreeing with the VLM's most confident read.

Three labelled boxes on M-8a

M-8a fully labelled: 244, 17, 336. The sequence panel reads as the photograph does.

What Changed

The benchmark re-scoped the pipeline.

  1. Segment the inscription band into sign boxes. This is the open problem: the signs abut, and projection heuristics fail. The labelling tool's boxes are the training data
  2. Classify each box locally against the 417 templates and return a shortlist
  3. Ask the VLM to pick from the shortlist, a 5-way choice instead of a 417-way one