02 / Mnemos - longitudinal evaluation
Evaluating Memory Over Time
A synthetic persona moves through conversations, tools, files, apps, screens, and memory revisions. Evaluation questions arrive only after the evidence they need.
Conversation & tools
Filesystem & activity
Mnemos memory
Evaluation checkpoints
Independent event
Sep 02, 2026
Filesystem at this point
Mnemos state at this point
Playback paused · evaluation checkpoint
Expected answer
Evidence
FR fact recall
MU memory updating
TR temporal reasoning
CR conflict resolution
SF selective forgetting
PU personalization / user modeling
How HippoCamp-Chat-S is synthesized
First assemble a grounded evidence core. Only then expand it into a realistic, noisy history.
Persona workspace / synthesis field
Maya47 files
profile.md
relocation-notes.pdf
calendar-history.json
preferences.yaml
golden file slice
grounded facts
+ evolving facts selected across modalities
+ evolving facts selected across modalities
Chat S-12introduces the relocation plan
Chat S-31uses a file through a tool
Chat S-48updates the earlier fact
EEvidence field
Qeval question
Areference answer
unrelated
trip chat
trip chat
ambient
tool use
tool use
old
preference
preference
nearby
session
session
file
mention
mention
routine
chat
chat
Evidence core / stages 1-5
Select grounded or evolving facts; synthesize multi-session chat/tool trajectories; create Q, reference A, and gold evidence; validate grounding, chronology, answerability, sufficiency, and no leakage.
← reject: return to trajectory + evidence synthesis
Distractor expansion / stages 6-7
Expand only after review. Distractors cannot answer or contradict gold evidence; deliberate updates must be tagged and reflected in A.
Evaluate / stage 8
Assemble checkpoints; run a fixed agent/model.