KuhaData

Pull the data out of unstructured documents without typing it all by hand.

Role
Design and build, solo · Undergraduate thesis, UP Mindanao
Adviser
Dr. May Anne Mata
Built
MVP February 2025 · Thesis January–May 2026
Stack
Svelte 5, FastAPI, PaddleOCR, SQLite, Llama 3.1 8B
F1 score (Textract: 0.85)
0.96
Documents tested
300
Office staff tested it
11
Ease of use
4.31/5
KuhaData reviewing a scanned fax cover sheet: the scan on the left, and on the right the extracted recipient, sender, date and CC, each labelled with how it was found and how sure the system is.
Reviewing a scanned fax. Each field shows how it was found and how sure KuhaData is.

Why I built it

My thesis adviser came to me with a problem: office staff were buried in MOAs and contracts, copying the same details out of each one by hand.

I said I could fix it in a week.

The first version worked, and showed me how much more there was to get right, so it became my thesis. Like GovForms, it's automation for work nobody should have to do by hand.

The problem

The tools that do this well charge per page and send every document to the cloud. The free ones are built for developers, not for the staff who handle the documents.

Per page, Textract form extraction
$0.065
A month, Unstract's starter plan
≈₱23,500

Rules first, AI where it's needed

I wanted the extraction to be predictable and explainable, so rules come first. Rules alone missed too much, so each document goes through three steps:

  1. OCR reads the page, even upside down or skewed.

  2. Rules find each field by its label and score how sure they are.

  3. Anything under 85% sure, and every table, goes to a language model.

Designed to run locally; tested through Groq with the same model.

The processing queue with three scanned faxes: one completed, one at the OCR stage at 20%, and one queued.
Documents move through each step in the queue.

Templates anyone can set up

Staff describe what they need once: the fields, their formats, other names a label goes by, and a hint for the AI if a field is tricky. Tables too.

The template editor for a progress report: a Subject field set to Short Text, then a Distribution table whose first column is Account name, each with optional alternate names and a hint for the AI.
A template for a progress report, with a Distribution table.

Showing its work

Every value shows how it was found and how sure the system is. Anything under 85% is flagged for a person to check before it's approved, and everything exports to CSV.

Reviewing a response code request confirmation: the Recipient field is highlighted in yellow as low confidence, while Sender, Date and Brand are marked verified.
A low-confidence field, flagged in yellow for review.
The progress reports processed with one template, in a table with their subjects and account names, and an Export CSV button.
Every document processed with one template, ready to export.

What the numbers said

I tested it on 300 documents, 100 public sample forms and 200 real documents from Davao City offices, against Amazon Textract.

F1 on sample forms (Textract: 0.85)
0.96
F1 on Davao City documents (Textract: 0.87)
0.96
Precision on forms
99.4%

It beat Textract on every measure except recall on the sample forms: Textract picked up more of the text, but put less of it in the right field. The language model does most of the finding. On forms, the rules raised precision from 93.6% to 99.4%.

Where it's slow

Reading the page is the slow part: OCR takes 81–93% of the time on a laptop with no GPU. Documents over 50 pages are still too long.

One-page form
~14 s
Multi-page PDF
~53 s

Testing with real people

Eleven administrative staff from Davao City offices ran their own documents through it.

Usefulness
4.26/5
Ease of use
4.31/5

Building it

  1. Feb 2025 First version, in about a week
  2. Jan 2026 OCR and the rule engine
  3. Apr 2026 Full pipeline, AI step and review screen
  4. May 2026 PDFs, tables and testing
  5. Now Defended; talking with offices about a rollout

What I learned

  1. A one-week fix is rarely the whole job. The first version worked; making it reliable on real documents took a thesis.

  2. Measure before believing. I expected the rules to carry the system; testing showed where they actually help.

  3. Build for the person at the desk. Staff who weren't technical had a few questions, then rated it easy to use.