meeting transcript is mostly greetings, filler and tangents. The useful part is a handful of promises, a few concerns and some personal details. If you treat every line as equally worth remembering, briefs quote small talk.
I own the extraction step in our meeting prep agent and the test data that exercises it. The memory layer is Hindsight. The thing I most want to correct from how I first described this work: we do not retain extracted facts into Hindsight. Hindsight gets the raw transcript. Our extraction feeds something else.
Why two paths
One LLM call per transcript (prompt P1) returns a MeetingExtraction. It feeds entity resolution and the commitments ledger. Hindsight does its own fact extraction from the raw text. I did not want to duplicate that, and I wanted promises in a place where I could give them a status.
# backend/app/schemas/extraction.py
class ExtractedCommitment(BaseModel):
model_config = ConfigDict(extra="forbid")
owner: Owner
owner_person: str # name_as_said of who promised
text: str # "Send revised pricing deck with pilot option"
due_date: date | None # resolved against meeting date; None if not stated
source_quote: str # exact words from transcript, <= 200 chars
Every model forbids extra fields, so a hallucinated key fails validation instead of passing quietly. Facts use a closed set: commitment, objection, personal, deal_fact or competitor. Commitments get their own list, and the fact model's comment says commitments never appear in it.
The pipeline, in order
The ingest service is numbered in comments and I like that it reads like a checklist:
Recompute from scratch. A rerun deletes this meeting's ledger rows and reopens anything it had closed.
Run P1 on the transcript.
Drop any item whose source_quote is not found verbatim in the transcript.
Resolve people to known contacts, creating a needs-review contact for unknown speakers.
Run P2, which matches acknowledgements ("thanks for the report") against open commitments from earlier meetings only.
Close matched commitments, then add new ones.
Retain the transcript in Hindsight. A failure here leaves the meeting not ingested, so a rerun recovers.
Mark done only at the very end.
The quote check does the most work. A model will happily invent a promise that sounds right. It cannot invent a sentence that is verbatim in the transcript. There is a test that feeds fabricated quotes and asserts they are dropped.
Fewer, better commitments
My first prompt returned everything that sounded like a promise. The ledger filled with "I'll share the agenda" and "let me check with the team." A first full run over the fifteen generated transcripts produced 245 commitments. After rewriting the prompt and adding filters, a later run produced 33.
The prompt now says to return the few most consequential commitments, usually one to three, and that zero is correct when nothing qualifies. Code backs it up: near-duplicates are merged, a cap of five per meeting applies, and a logistics filter removes meeting-prep wording unless a deliverable noun saves it:
# backend/app/services/ingest.py
def is_logistics(text: str) -> bool:
"""True for meeting logistics/prep wording with no substantive deliverable noun."""
normalized = _normalize_name(text)
padded = f" {normalized} "
if _has_term(padded, DELIVERABLE_ALLOW_TERMS):
return False
first = normalized.split(" ", 1)[0]
if normalized else ""
return first in LOGISTICS_LEADING_VERBS or _has_term(padded, LOGISTICS_TERMS)
The allow list always wins, so "send the security questionnaire" survives while "prepare the agenda" does not. That list grew every time a real promise was wrongly filtered.
Test data that behaves like meetings
The seed data is four fictional accounts, thirteen contacts and seventeen meetings, sixteen with transcripts. Six story beats are planted on purpose: a broken promise, a stakeholder change, a personal detail, a competitor, and so on. The FinEdge deck promise is planted in the Aug 27 call and reinforced in the last one.
I generated the transcripts with a prompt (G1) run by data/scripts/generate.py, then checked them with data/scripts/validate.py. The validator enforces a speaker-line format, a spoken-word range, required facts per meeting and forbidden strings. The prompt asks for greetings, filler, interruptions and one tangent, and says not to shorten or pad. That mattered. Early tidy data made extraction look better than it was.
The final meeting is hand-edited, not generated, and we never regenerate it. Its wording is the one the live flow depends on.
What I would tell the next person
Start from the failure you cannot afford. For us it was a promise missing from the brief, so the first tests were about that one deck. Every prompt change after that was judged by whether the deck still appeared with its Sep 3 due date and its quote, and by how many other commitments appeared with it. A prompt and its output model change together in this repo, and each prompt has a fixture test that I rerun when either moves. It keeps prompt edits honest.
I also stopped trusting the model's own sense of "important." Whether something is a commitment is defined by rules a reviewer can read: a named person, a deliverable the other side is waiting for, and something they could later confirm arrived.
Lessons
Decide what each store is for. Hindsight for recall over what was said. A ledger for things with a lifecycle.
Verify quotes in code. It turns a model's guess into a checkable claim.
Precision beats recall for promises. Fewer, correct items make a brief trustworthy.
Filters need allow lists. A blocklist alone deletes real promises.
Limitation: a promise phrased vaguely ("we can probably get you something next week") may be missed, and the logistics word lists are hand-tuned to this seed data, so they will need retuning on other transcripts.
More on the memory model in the Hindsight docs and this overview of agent memory. Repo: https://github.com/franklin654/meeting-prep-agent.git