Web pages in. Typed records out.
moocher finds unstructured information across the web, structures it into schema.org-typed records — 20+ types — and enhances them against reference datasets, public records, and map data. Confidence-scored, and legible to the machines that now answer the questions.
- name:
- startDate:
- location:
messy page in, labeled record out.
The facts exist. The records don't.
Across the public web, the facts are all out there — hours, dates, addresses, menus — but published as presentation, not data: HTML arranged for human eyes, shaped differently on every site that carries it. No shared types, no consistent fields, no key to join on. People copy-paste; machines guess.
The same trap exists inside organizations. An aging site or CMS holds years of records with no export path — the database undocumented, the vendor long gone — and the rendered pages become the only surviving source of truth. Rebuilding means re-keying it all by hand, or walking away from the data.
Find. Structure. Enhance.
Three stages, in order. Each does the simplest thing that works — and escalates only when the page demands it.
Find
Point moocher at a listing page and it identifies the entities, then follows each to its own detail page — autonomous discovery, not a URL list you maintain. Given no more than a name and a town, it can turn up an entity's official site on its own; when a site resists, fetching escalates instead of failing.
Structure
Every record comes back typed against schema.org — Event, Recipe, Person, LocalBusiness, 20+ types in all — and the type is detected, not configured: hand it a page cold and it works out what it holds before labeling a single field. Every record carries a confidence score; the system reports its own uncertainty instead of hiding it.
Enhance
A single page is one witness. moocher checks each record against reference datasets, public records, and map data, keeping field-level source attribution as it goes — every fact in the finished record traces back to where it came from. One record, several witnesses.
Structured means schema.org.
Schema.org is a shared vocabulary for describing the things on a web page — an Event, a Recipe, a LocalBusiness — created by the major search engines so machines could agree on what a thing is. It's the closest thing the public web has to a common data model.
Two kinds of machines read those labels. Search engines use them to decide what to show; AI assistants and answer engines use them to decide what to say — a typed, unambiguous fact is something a retrieval pipeline can lift straight into an answer, no parsing step to get wrong.
Unlabeled content gets interpreted — a generous word for guessed at. Guesses drop fields, mangle dates, skip entities. Labeled data is read as written: the difference between being quoted and being paraphrased.
Every record moocher hands back arrives already labeled — schema.org types, schema.org property names — ready to publish as structured data or feed whatever sits downstream. It speaks 20+ types natively and knows the deep end of the vocabulary: it can tell a ComedyEvent from a TheaterEvent.
- Event
- Recipe
- LocalBusiness
- Person
- Menu
- Offer
The engineering, briefly.
Escalation, not brute force
Structured data is read directly where a site publishes it; learned patterns cover pages the system has seen; model inference is reserved for pages that genuinely need it. The best model call is the one you don't make.
Browser-grade fetching
JavaScript rendering and dynamic content are table stakes. When a site pushes back, fetching escalates through progressively more capable strategies instead of giving up at the first refusal.
Native schema.org typing
Records typed against 20+ schema.org types — EventProductRecipeArticlePersonLocalBusiness — with the type auto-detected from the page, not configured per site.
Confidence-scored output
Every record carries a confidence score and a completeness measure. A system that reports its own uncertainty is one you can put review gates on — silent failure is the alternative.
Autonomous discovery
Start from one listing page; it finds the entities and follows each to its own page. Given a name and a town, it can locate the official site itself.
Multi-source enrichment
Finished records are checked against reference datasets, public records, and map data — field-level attribution included, so every fact traces to the source that supplied it.
What you could point it at.
Every show, market, and one-off in a scene, announced across unrelated sites in unrelated formats — coming back as Event records: name, startDate, location, the same fields every time.
Every business along the district you care about — what it is, where it is, when it's open — as LocalBusiness records instead of a browser session and a spreadsheet nobody maintains.
Recipe pages that bury the ingredients under an essay; menus published for reading, not for use — returned as Recipe and Menu records: ingredients, steps, items, each in its own field.
The small outlets and blogs covering the beat you follow — captured as Article records: headline, author, date, every source in the same shape.
The export the old system never gave you.
An aging site holds years of records with no export path. The database is undocumented, the vendor is gone, and the rendered pages are the last complete copy of the data. The usual way out is re-keying it all by hand.
Hand moocher the old site's sitemap — or a plain list of URLs in a text or CSV file — and it reads every page. What it finds is mapped into one common shape — spreadsheet, JSON, CSV, schema.org-typed records underneath — and cleaned in transit: spelling corrected, formatting normalized, inconsistencies reconciled.
What comes back is ready to import into the system that comes next. The records cross over, the old system retires, and nobody re-keys years of data by hand.
Describe the system you're replacing — and see what comes back.

The dataset you want doesn't exist yet.
It's out there — scattered, unlabeled, inconsistent: exactly the input moocher was built to take. There's no signup form and no sales funnel; there's Bo — one person who built moocher, runs it, and answers the email. Describe the data you wish you had — and see what comes back.
Email Bo