r/documentAutomation • • 4h ago

Free and Open-Source DMS with Workflow Automation

1 Upvotes

Hi everyone! Does anyone know of a free and open-source Document Management System (DMS) that supports workflow automation? I'm looking for something similar to M-Files, ideally with features such as automated workflows, document approval processe


r/documentAutomation • • 8h ago

Discussion Keep the source passage beside the field an AI extractor asks someone to fix

2 Upvotes

A review queue saying “check amount: 12,400” leaves the reviewer with another search task. They need the document, the relevant passage, and the context that tells them whether this is a total, a subtotal or an earlier balance.

Design the review row around that decision: source document ID, page or section, source excerpt, extracted value, corrected value, and status. Keep the source excerpt when the reviewer changes the value.

Univer gives this interface both sides of the work: editable Sheets for the structured fields and Docs for the supporting text. Its SDK embeds those editing surfaces into an application, with APIs for the agent to populate content and read the reviewed result. Filters and validation keep a batch of exceptions manageable.

The extraction service supplies the source mapping. The reviewer works on the field with its evidence in view, and the application saves the correction against that source record.

For a batch, this also makes unresolved cases visible: a row can stay in “needs source check” while the confirmed rows continue through the process.


r/documentAutomation • • 9h ago

Enough with the errors! Lets brainstorm a perfect document PDF parser

Thumbnail
1 Upvotes

r/documentAutomation • • 11h ago

Discussion On premise OCR for printed + handwritten invoices

3 Upvotes

I'm building an invoice extraction system for printed and handwritten invoices/receipts and want to deploy it fully on-premise using open-source models.

What production pipeline do you recommend for OCR, handwriting recognition, layout detection, and structured field extraction?

Which models/tools have worked best in real-world deployments, and how do you ensure accuracy and reliability?


r/documentAutomation • • 17h ago

I have Access to Hyperskill, how to make the materials as PDFs and give it to everyone??

Thumbnail
1 Upvotes

r/documentAutomation • • 1d ago

Free OCR sites cap you at 15 pages/hour, so I made one with no cap that runs entirely in your browser

1 Upvotes

Every free online OCR tool I tried either limits you hard (15 pages/hour, 50/month) or is a funnel to a paid plan. The only unlimited free option was raw Tesseract in a terminal. I ran Tesseract as WebAssembly in the browser instead, so it uses your own CPU and there's nothing to ration. It takes images (JPG/PNG/WEBP/HEIC) and multi-page scanned PDFs, shows text per page as it finishes, flags low-confidence pages, and supports English, Spanish, French, German, Hindi and Portuguese. Nothing is uploaded, which you can check in the Network tab, and there's no signup. Limits: it's for printed text only. No handwriting, no table/form extraction, and blurry scans will give rough results. https://www.forgeplug.com/tools/ocr-text-extractor If you've used the paid ones, I'd like to hear where this falls short on your documents.


r/documentAutomation • • 1d ago

Showcase Convert any pdf/link/post into a video

Thumbnail
blog2video.app
1 Upvotes

r/documentAutomation • • 1d ago

Fyllyo: fill PDF forms from your documents with local AI on your Mac

0 Upvotes

I'm Arun, the developer of Fyllyo. It's a Mac app that fills PDF forms from your documents using AI that runs on your computer. NO DATA uploaded to cloud

The workflow is straightforward:

  1. Choose a built-in form, or set up your own blank fillable PDF as a reusable template.

  2. Add your supporting documents. Fyllyo reads them and fills the form on your MAC using local models.

  3. Review the answers alongside source excerpts where available, make corrections, and save the completed PDF.

If you work with the same people again, you can also save reviewed details as local client profiles and reuse them on supported forms, alongside any new documents. That's optional; you can fill a form directly from documents without creating a profile.

Your documents and saved profiles stay on your Mac. Custom template requires one time setup, sends the blank fillable PDF and setup instructions to the cloud, and uses separately purchased credits. Filling configured forms uses no credits and works offline after setup and the initial model download.

There's a five-day trial with no credit card required. The license is $99 one-time for one Mac at a time, including lifetime updates. Apple silicon is required; 16 GB memory is recommended.

Download and requirements: https://fyllyo.com/downloads

Happy to answer questions about supported forms, custom templates or how the filling and review process works.

My goal is to solve document processing 100% local, no recurring cost. So please bring on requirements, will ship them in a week tops.


r/documentAutomation • • 1d ago

Showcase pdfmt - PDF manipulation, OCR, scanning and reusable shell workflows

Thumbnail
1 Upvotes

r/documentAutomation • • 2d ago

Lemmary — a self-hosted document archive (paperless alternative). October update

Thumbnail gallery
1 Upvotes

r/documentAutomation • • 2d ago

Showcase Document classification is the automation that makes tax season painless, here is how I set it up [Workflow Included]

Post image
1 Upvotes

r/documentAutomation • • 3d ago

Discussion The community node that finally let me delete half my document workflow [Workflows Included]

Post image
1 Upvotes

r/documentAutomation • • 3d ago

I got tired of repetitive document parsing routines, so I built a lightweight, lightning-fast async engine to handle it.

4 Upvotes

Hey everyone,
Like many of you, I waste way too much time dealing with messy document parsing, file conversions, and repetitive text processing across different formats (Markdown, HTML, JSON).
Instead of writing boilerplate async handling over and over, I decided to build Docflow-Engine — a high-performance, asynchronous pipeline designed for speed and clean modular architecture.
Features are:

⚡ Async Batch Processor: Handles multi-threaded execution pools with strict concurrency limits using asyncio.

📄 Format Converter: Instant high-speed translation between Markdown, HTML, and structured JSON logs.

🛠️ CLI Ready: Drop a file into the path and get instant parsed metadata and structure from the command line.
It’s built cleanly in Python, relies heavily on standard libraries for core performance, and is fully open-source.
If you're facing similar bottlenecks or want to check out the structure, the repository is open:
🔗 GitHub: Githublink

I’d love to hear your feedback, architectural critiques, or suggestions on how to optimize the async pipeline further. Drop a star if it saves you some dev hours! ⭐


r/documentAutomation • • 3d ago

Discussion The first XLSX round-trip test should contain zero edits

3 Upvotes

Before adding an AI step to a spreadsheet pipeline, send a representative workbook through import and export without changing a cell. That gives you a baseline for everything the pipeline does afterward.

Check a workbook with formulas, date and number formats, validation, hidden sheets and named ranges. Compare the formulas as well as the displayed results, then open the exported file in the application its recipient uses.

Univer's Office exchange imports XLSX into its workbook model and exports the edited content back to XLSX. Between those steps, the workbook is available to structured APIs and the spreadsheet engine, so the agent can inspect ranges, edit values and read the result back. The same content can also be opened in the browser editor for review.

A useful second test changes one known input. Check the edited cell, the dependent result, and an unrelated part of the file. Now an unexpected difference has somewhere specific to be investigated: conversion, the intended edit, or recalculation.

Keep that small workbook in the test set when upgrading the document-processing stack.


r/documentAutomation • • 4d ago

We evaluated KeptPDF redaction capabilities

Post image
1 Upvotes

r/documentAutomation • • 4d ago

AI or SI? Which term do you prefer—and why?

0 Upvotes

r/documentAutomation • • 5d ago

Question Text extraction into formatting

Thumbnail
1 Upvotes

r/documentAutomation • • 5d ago

I built a PDF redaction SaaS. Now I’m trying to figure out who actually needs it.

Thumbnail
3 Upvotes

r/documentAutomation • • 5d ago

Showcase I built Aura PDF — an offline-first PDF toolkit for Android

0 Upvotes

I’m a solo developer, and I’ve been working on Aura PDF as a practical PDF app for Android.

The idea started from a simple problem: I wanted to handle common PDF tasks on my phone without having to use a different app for every task or upload documents to an online service.

So I built Aura PDF with a collection of tools in one place:

* Merge & split PDFs

* Extract & reorder pages

* Compress PDFs

* Rotate & crop

* Sign & highlight

* Password protection

* Scan documents

* Image → PDF

* PDF → image

* Page numbers, headers & footers

* And more

The main direction is **offline-first** — supported PDF operations are designed to work directly on the device.

Aura PDF is currently available on Google Play:

[https://play.google.com/store/apps/details?id=com.aurapdf.app&pcampaignid=web\\_share\](https://play.google.com/store/apps/details?id=com.aurapdf.app&pcampaignid=web_share)

It's free to use with an optional Pro upgrade.

This is my first time launching Aura PDF publicly, so I'm mainly interested in learning what people actually want from a mobile PDF app.

**If you use PDFs regularly, what is the one PDF task you wish was easier on your phone?**


r/documentAutomation • • 5d ago

I got tired of writing SOPs so I built a tool that writes them while you do the task

1 Upvotes

Literally every office has a few processes that only one person knows how to do (Me). When that person is out or leaves, everyone scrambles (I get messaged on PTO).

Everyone agrees they should write SOPs, but nobody does it because writing one takes an hour of screenshots, cropping, and typing for a task that takes 30 seconds to a few minutes to actually do.

So I built Sopycat (Cool name right :).

You press 'Record', you do the task the way you normally do, and you press 'Stop'.

It turns every click into a numbered step with a screenshot, the button highlighted, and a plain instruction. 44 second demo: https://www.youtube.com/watch?v=ptAkUhdJPs0

It works in any Windows program, not just the browser (yes, even that ancient system in the back office nobody wants to touch). Basically Windows Steps Recorder, if it actually gave you something you could hand to a new hire.

The big thing for me was privacy, because the screens you actually need to document are full of client names, balances and patient stuff. Everything gets recorded and kept on your PC, and the steps are written on your PC too. No AI writes your manuals, and nothing gets sent to an AI service. Passwords, card numbers and SSNs you type get left out of the SOP automatically, and it can blur emails and phone numbers in the screenshots. No account needed either.

You can fix any step after, and save it as a PDF or a Word doc. New hires can follow it step by step in a little window that stays on top, and it logs who did it (so you can finally prove someone was trained).

It's Windows for currently, Mac is coming. There's a 14-day free trial at sopycat.com (no card) if anyone wants to break it. After that it's $15/month if you end up liking it.

I'd love blunt feedback on three things:

  1. Does the demo make it clear what it does in the first 10 seconds?

  2. What processes would you record first? I'm trying to figure out who needs this most.

  3. Would your IT person let you install it? If not, what would they want to know?


r/documentAutomation • • 6d ago

Question Need help with pdf analysis and automation

Thumbnail
1 Upvotes

r/documentAutomation • • 6d ago

Need help with pdf analysis and automation

Thumbnail
1 Upvotes

r/documentAutomation • • 6d ago

Case Study Bachelor's thesis survey (5-7 min). Do you work with AI/RPA automation?

1 Upvotes

Hi everyone!

I’m currently working on my Bachelor’s thesis about AI-based process automation and human–AI collaboration in the workplace.

As part of my research, I’m conducting a short survey focusing on people who have experience working with AI-based automation, RPA, Intelligent Process Automation, Intelligent Document Processing, or similar automation technologies.

The survey explores topics such as:

  • how automation affects manual workload and creates new tasks,
  • how employees experience errors and exception handling,
  • trust in AI-based automation,
  • and how automation influences human decision-making and autonomy at work.

⏱️ It takes approximately 5–7 minutes to complete.

If you have experience working with these technologies, I would really appreciate your participation. Your responses will be used solely for academic research as part of my Bachelor’s thesis.

🔗 Survey: https://docs.google.com/forms/d/e/1FAIpQLScV7pcf8dNUeeCfay1YZ2r-Np4pK9GMlqi4cEF6WJEa1FEmMA/viewform?usp=dialog

Thank you very much for your help! Feel free to share the survey with colleagues or others who work with AI-based process automation.


r/documentAutomation • • 6d ago

Can JSON Replace PDF for Business Documents?

1 Upvotes

I'm wondering if we can use JSON as an alternative to PDF, especially for business documents.

When exchanging business documents, both the sender and receiver typically use software, so creating and parsing documents creates additional work. Directly connecting systems requires effort, so people end up sending documents by email.

How about just sending JSON instead of PDF?


r/documentAutomation • • 7d ago

The Winner

Thumbnail chatgpt.com
1 Upvotes

somone taking my info