Define the test before sending a message

Record the app, plan, interface, character, mode, visible model information, memory settings and manual notes. Fix the prompts, answer key and scoring rules in advance. Keep automatic conversation-derived memory and manually supplied notes in different test conditions.

Use fictional adult profiles and neutral details. Repeated runs in one established account may share history; label them repeated trials rather than claiming independent accounts. Do not create accounts in ways that violate a service’s rules.

Assign separate facts to each checkpoint

Prepare three matched sets with different values. Set A belongs to the 24-hour check, B to day 7 and C to day 30. Each set contains a stable preference, an explicitly corrected schedule and a never-supplied detail. Introduce the facts and corrections in the same order for each service, and timestamp them.

Measure the checkpoint delay from the final seed or correction for that set. Write down the actual interval. A late checkpoint should be reported as late, not relabeled as an exact interval. Do not preview the answer to a later set during an earlier check.

Run three separately initialized trials per condition where the service permits that setup. This is an exploratory sample, not enough to establish population-wide reliability. Record reuse of accounts or characters as a limitation.

Visual field note / 08Separate facts. Real waiting periods.. Measure each interval from that set's final seed or correction.
Visual summary of the guide. Examples are fictional; this is not a product result. Open full-size diagram (opens in a new tab)

Keep the question free of its answer

Ask “When is the club currently scheduled?” rather than “Do you remember it moved to Tuesday?” Capture the first complete reply. If you later provide a hint or regenerate, save that as a repair attempt with its own record; never overwrite the initial observation.

For unknown details, explicitly ask for established information, not creative invention. For a new-chat condition, check that the claimed memory scope includes it. Do not assume closing an app removes earlier context.

Use categorical outcomes before aggregates

Outcome rules for a single question
LabelRule
CorrectSupplies all required information without a material contradiction
PartialSupplies some required information, omitting the rest without contradiction
MissCannot supply a known fact or gives an incorrect answer
Appropriate uncertaintyRecognizes that the requested detail was never established
Invented detailPresents an unsupported detail as established history; flag alongside the main outcome
Unavailable / invalidUnsupported condition or a capture/setup failure; explain and exclude from valid counts

Report correct known-fact answers out of valid known-fact questions separately from appropriate uncertainty on unknown-detail questions. Keep partial answers visible. Missing or unavailable observations are not zero scores. Preserve any reviewer disagreement rather than silently averaging it away.

Publish the limits with the evidence

Attach the prompt set, original statements, correction timestamps, complete replies, assistance log and configuration. Redact personal information. An observed answer does not reveal a database design, prove deletion or establish emotional understanding.

LongMemEval provides research categories including updates, time and abstention. This small consumer protocol is an editorial adaptation, not that benchmark and not a reproduction of its results.

Download the blank observation log. Current evidence status: no independent product results; no winner declared.