Multi-agent metadata curation
Manual curation cannot keep pace and automated curation misses fields. This one is measured on recall across all 23.
Details
Metadata for a public dataset is scattered across the entry, the paper, and the supplements. A curator reads all three.
First author. I own the agent architecture, the field extraction, and the recall measurement.
- An orchestrator delegates retrieval, parsing, ontology mapping, and inference to expert sub-agents.
- Recall is measured across all 23 fields, including original terms, normalized terms, and ontology identifiers.
- The system must scale to hundreds of fields without losing precision.
- Preprint on bioRxiv, DOI 10.1101/2025.06.10.658658.
- 93% average recall across 23 fields, reported in the abstract.