KuhaData
Pull the data out of unstructured documents without typing it all by hand.
- F1 score (Textract: 0.85)
- 0.96
- Documents tested
- 300
- Office staff tested it
- 11
- Ease of use
- 4.31/5

Why I built it
My thesis adviser came to me with a problem: office staff were buried in MOAs and contracts, copying the same details out of each one by hand.
I said I could fix it in a week.
The first version worked, and showed me how much more there was to get right, so it became my thesis. Like GovForms, it's automation for work nobody should have to do by hand.
The problem
The tools that do this well charge per page and send every document to the cloud. The free ones are built for developers, not for the staff who handle the documents.
- Per page, Textract form extraction
- $0.065
- A month, Unstract's starter plan
- ≈₱23,500
Rules first, AI where it's needed
I wanted the extraction to be predictable and explainable, so rules come first. Rules alone missed too much, so each document goes through three steps:
-
OCR reads the page, even upside down or skewed.
-
Rules find each field by its label and score how sure they are.
-
Anything under 85% sure, and every table, goes to a language model.
Designed to run locally; tested through Groq with the same model.

Templates anyone can set up
Staff describe what they need once: the fields, their formats, other names a label goes by, and a hint for the AI if a field is tricky. Tables too.

Showing its work
Every value shows how it was found and how sure the system is. Anything under 85% is flagged for a person to check before it's approved, and everything exports to CSV.


What the numbers said
I tested it on 300 documents, 100 public sample forms and 200 real documents from Davao City offices, against Amazon Textract.
- F1 on sample forms (Textract: 0.85)
- 0.96
- F1 on Davao City documents (Textract: 0.87)
- 0.96
- Precision on forms
- 99.4%
It beat Textract on every measure except recall on the sample forms: Textract picked up more of the text, but put less of it in the right field. The language model does most of the finding. On forms, the rules raised precision from 93.6% to 99.4%.
Where it's slow
Reading the page is the slow part: OCR takes 81–93% of the time on a laptop with no GPU. Documents over 50 pages are still too long.
- One-page form
- ~14 s
- Multi-page PDF
- ~53 s
Testing with real people
Eleven administrative staff from Davao City offices ran their own documents through it.
- Usefulness
- 4.26/5
- Ease of use
- 4.31/5
Building it
- Feb 2025 First version, in about a week
- Jan 2026 OCR and the rule engine
- Apr 2026 Full pipeline, AI step and review screen
- May 2026 PDFs, tables and testing
- Now Defended; talking with offices about a rollout
What I learned
-
A one-week fix is rarely the whole job. The first version worked; making it reliable on real documents took a thesis.
-
Measure before believing. I expected the rules to carry the system; testing showed where they actually help.
-
Build for the person at the desk. Staff who weren't technical had a few questions, then rated it easy to use.