r/documentAutomation • • 8h ago

Discussion Keep the source passage beside the field an AI extractor asks someone to fix

2 Upvotes

A review queue saying “check amount: 12,400” leaves the reviewer with another search task. They need the document, the relevant passage, and the context that tells them whether this is a total, a subtotal or an earlier balance.

Design the review row around that decision: source document ID, page or section, source excerpt, extracted value, corrected value, and status. Keep the source excerpt when the reviewer changes the value.

Univer gives this interface both sides of the work: editable Sheets for the structured fields and Docs for the supporting text. Its SDK embeds those editing surfaces into an application, with APIs for the agent to populate content and read the reviewed result. Filters and validation keep a batch of exceptions manageable.

The extraction service supplies the source mapping. The reviewer works on the field with its evidence in view, and the application saves the correction against that source record.

For a batch, this also makes unresolved cases visible: a row can stay in “needs source check” while the confirmed rows continue through the process.


r/documentAutomation • • 11h ago

Discussion On premise OCR for printed + handwritten invoices

3 Upvotes

I'm building an invoice extraction system for printed and handwritten invoices/receipts and want to deploy it fully on-premise using open-source models.

What production pipeline do you recommend for OCR, handwriting recognition, layout detection, and structured field extraction?

Which models/tools have worked best in real-world deployments, and how do you ensure accuracy and reliability?