Case study — Scripps Institution of Oceanography
Search a cruise report instead of reading it.
Very long research cruise reports, read and indexed, so somebody after one detail can search for it and get an answer they can check — instead of reading end to end and hoping they remember the right page.
In their words
“During his time with CCHDO group at Scripps Institution of Oceanography, Noah developed an AI-assisted metadata extraction workflow for scientific cruise reports and helped move it from prototype development toward production use. His work was validated against manual extraction and demonstrated that the workflow could process our report reliably while substantially reducing the amount of manual effort required.”
How it works
The problem
A cruise report is the record of one voyage — what the ship did, what was deployed and when, what went wrong and what was done about it. It gets written once, carefully, and then it sits there. Later somebody needs one specific thing out of it, and keyword search only works if you already guessed the wording the author used.
What it does
It reads each report — including the parts that are really tables pretending to be prose — and breaks it into passages small enough to be an answer and large enough to still make sense alone. Those go into an index built for retrieval rather than keyword matching, so a question asked in ordinary language finds the passage that answers it.
What comes back
The passage, and which report and section it came from. That second half is the point. An answer you can't trace back to a page is a claim, not a citation, and this reader is a scientist.
How it was checked
Against manual extraction — the same reports done by hand, compared with what the workflow returned. Confidently wrong is worse than nothing: nothing sends you back to the document, and wrong doesn’t.
It is containerized, like everything we build, so it runs wherever it is put — an institution's own machines, or infrastructure we run for you. The documents and the index live together wherever that is, and neither is handed to a third party for the search to work.
Six things this page doesn’t claim
The quote above is somebody else’s account of the work, and we haven’t put numbers next to it that we can’t confirm against the record. A figure we rounded up would undo the point of the rest of the site.
- How many reports are in the index, and roughly how many pages that comes to.
- How long it used to take to find something, measured rather than remembered.
- The accuracy figure from that validation, and how big the hand-checked set was.
- When the work happened and over what period.
- Which roles use it day to day.
- Whether it runs on hardware at the institution or on a machine locally.
If you want those before you’d take this seriously, that’s the right instinct. Ask, and you’ll get whatever is confirmed at the time you ask.
Send this to someone
Long documents are long documents.
Cruise reports are unusual. The shape of the problem is not. Environmental impact reports, geotechnical reports, permits, contracts, monitoring records and hearing transcripts are all written once, read under pressure, and searched badly. Tell us what yours are and what people keep needing out of them — twenty minutes is usually enough to work out whether the same approach fits.
Omnentis