The four output goals — what your captures unlock
Each Heirloom capture produces multiple outputs. The more you capture, the more outputs become available. Here's exactly how much you need for each one.
Audio archive (the foundation)Basic
Every memoir starts with audio. Each session can be a single 15-minute call answering 3–5 guided questions. Quality requirements:
- Recorded in a quiet room (no fan, no TV in the background)
- Phone microphone is fine — no studio needed
- WAV or MP3, 44.1 kHz mono is enough
- Single speaker (you can both speak — we identify and separate)
- WhatsApp voice notes work perfectly for daily capture
Voice clone (recognizable likeness)Recognizable
The voice clone needs clean audio without background music, multiple speakers, or phone-line distortion. From any 30-minute clean session you'll get a recognizable clone. The richer the audio, the more natural the clone — including emotion, laughter, and natural pauses.
- 30 min = recognizable but slightly synthetic
- 60 min = clearly them, mostly natural
- 2 hr = indistinguishable from a real recording
- 6 hr+ = full emotional range; whispers, laughter, tears
Photo archiveVisual memoir
Photos serve two jobs: the print-ready Memory Book and the family archive. Quality requirements:
- Minimum 10: 4 childhood, 3 young adult, 3 elder. Faces visible, well-lit, front-facing.
- For the Memory Book: 20–30 photos your family curates and places in the chapters they choose
- Phone photos OK; old paper photos scanned at 300 DPI minimum
- Avoid: tiny resolution thumbnails, group photos where face isn't visible, sunglasses or shadow on face
Live video (most powerful, optional)Best output
Live video of the subject answering questions is far more powerful than animated still photos. If they're alive and willing, capture short video clips of them telling key stories. Use cases:
- 30-second clips for each life chapter (childhood, marriage, career, parenthood)
- 1-minute clip per major story (how they met spouse, day they emigrated, war story)
- Phone camera in landscape, eye-level, 4K if possible (1080p is fine)
- Good lighting (window light is enough)
- Kept in your family archive alongside the Memory Book and the interactive AI experience
Documents + ephemeraBonus
Letters, marriage certificates, immigration papers, military discharge papers, recipe cards in their handwriting, tickets from key events. Scan or photograph. Your family can weave these into the Memory Book chapters as you review and edit them.
The auto-processing pipeline (what Heirloom does behind the scenes)
- Ingest: WhatsApp/phone-call/app-recorded audio uploads to encrypted storage. EXIF data stripped. Each file tagged with memoir_id + session_id.
- Transcribe: Deepgram processes audio → text with speaker diarization (separating you from the subject). Transcript stored per session.
- Translate (if needed): Source language preserved; English translation generated for descendants who don't speak the original.
- Theme extract: Anthropic Claude reads the transcript, identifies which question(s) it answered, tags themes (childhood, marriage, immigration, etc.).
- Voice model train: Once 30+ minutes of clean audio accumulates, ElevenLabs voice-clone build is queued (consent gate must be PASSED).
- Photo enrich: Photos are face-detected, age-estimated, OCR'd for text in the image. Auto-tagged for chapter alignment.
- Memoir compile: When the family compiles the Memory Book, AI assembles draft chapters using only source-linked content (no hallucination). Your family reviews and edits every chapter, then downloads the print-ready PDF.
- AI ancestor: Voice clone + transcript bank + photo gallery → conversational interface. Strictly source-linked. Says "I don't know" when asked something the subject never spoke about.
Every step logged for audit. Every transcript reviewable. Every model deletable on request.
For best results — the 90-minute starter recipe
If you want to give your loved one the best possible Heirloom experience, give us this in week one:
- 30 minutes of clean voice answering 4–5 of our starter questions
- 10 high-quality photos covering childhood, young adult, parent, elder ages
- One 60-second video clip of them speaking in their natural voice
- One scanned document with their handwriting on it (a recipe, a letter, anything)
That's a complete, production-ready foundation. Everything after week one builds on this.
What to avoid (so we don't have to ask twice)
- Audio recorded with TV/music in the background
- Photos with multiple people where the subject's face is unclear
- Voice memos where you talk over them
- Photos that are very low resolution (under 600px wide)
- Recordings while they have a cold (we recommend rescheduling — voice clone quality matters)