Where we startedManual review worked, until volume started winning
Every document had to be read and converted into structured data. The reviewers were good at it, but the queue grew whenever volume grew. Hiring more people would have bought time. It would not have changed the shape of the problem.
We could extract fields with an LLM, but an impressive demo was not enough. A wrong value could affect a downstream decision. Before we discussed scale, I wanted the team to answer a more basic question: when the system is wrong, how will we know, and who will act?
The first product decisionMake confidence useful
We stopped debating manual versus automated. Each field could take a different route. A high-confidence extraction could continue. A low-confidence or high-impact field went to a person. A failed step became a visible exception instead of disappearing in a pipeline.
That gave us a sensible 0 to 1 boundary. We did not need to solve every document type or automate every field. We needed one complete path that could extract, route, review and leave an audit trail. A narrow product running with real work taught us more than a broad prototype.
What the team builtThe model was only one part of the system
The production work was mostly around the model:
- Field-level confidence rules so one uncertain value did not make the whole document opaque.
- Review queues that showed the source, extraction and reason for escalation in one place.
- Visible exceptions with ownership, retries and enough context to investigate a failure.
- Traceability from the original document to the final value and every human correction in between.
We started with conservative thresholds. That meant more human review in the early weeks, which was fine. The team could see the failure patterns, improve prompts and rules, and loosen thresholds with evidence. Trust grew because the system was honest about uncertainty.
I did not approve straight-through automation followed by occasional sampling. That approach would find some mistakes after they had already moved downstream. For this product, review was part of the workflow, not a quality exercise performed later.
From first release to scaleChange the unit of human work
The first release proved that documents could move safely through the full path. Scaling it was a different job. We watched queue age, correction patterns, document mix and reviewer load. We improved the rules where the same exception appeared repeatedly. We also made sure operational ownership did not sit with the few people who built the pilot.
Human review moved from the default path to the exception path. Reviewers spent less time confirming obvious values and more time on cases that needed judgment. The system remained readable: we could explain what it extracted, why it escalated and who changed it.
What I learned
I would still design the human workflow before optimizing the model. The review queue was not a temporary safety net. It was part of the product and remained important as volume grew.
I would invest earlier in confidence calibration and reviewer analytics. Some early thresholds came from judgment because that was the fastest way to start. Once the product was live, those numbers needed to come from observed errors and correction data.
The useful question was not how many people we could remove. It was how much important work the same team could handle safely.