Research · Measurement plan
How we will measure x4me2 review recovery
The homepage assumption sandbox uses invented values. This page states what we intend to count before data collection begins, so later results cannot quietly change the definition of success. It is a public measurement plan, not an independently archived preregistration.
Published 10 August 2026 · Updated 22 August 2026
The domain
x4me2 helps people recover reviews they already wrote on platforms, prove authorship, and host them outside the original platform. Payments from AI systems are planned, not live. It is a League project, so this will be an internal product measurement rather than an independent evaluation. Raw counts, including failures, will be published here for others to inspect.
The unit
The main unit is one completed review recovery. A session counts as complete only when all four conditions hold:
- Need matched: a person identifies specific reviews they authored on a platform.
- Terms accepted: they authorize recovery and independent hosting.
- Checks complete: the authorship check passes.
- Handoff done: the review is live under the person’s control in a movable form.
The counters
- Attempts
- Recovery sessions started, including abandoned and failed ones.
- Completions
- Sessions reaching the live-under-control state defined above.
- Completion rate
- Completed recoveries divided by all started sessions.
- AI use per attempt
- Metered model usage across every automated call in the recovery process, averaged across all attempts.
- Checking cost per completion
- Compute cost plus human review time. The hourly rate for human review will be stated with the results.
- Total cost per completion
- AI use, authorship checks, human review, and operating cost, divided by completed recoveries.
- Manual comparison
- The time and cost required to reconstruct an equivalent, verifiable record without x4me2. The comparison procedure and hourly rate must be published before collection begins.
Still to decide before measurement begins
- The observation window and minimum number of attempts.
- The exact event that marks the start of a recovery session.
- The hourly rate used for human review and the manual comparison.
- The procedure for deciding whether a recovered record is equivalent to the manual version.
These decisions will be added as a dated amendment before the first measured batch. Until then, the plan is incomplete.
Community feedback
The open question right now is simpler: would recovering a review be useful, what would make the process feel unsafe, and which decisions must remain with the author? Reactions and objections are welcome through the contact form. When measurement begins, batches will include failures and abandoned sessions.
Results that would weaken the claim
- A completed recovery costs more than reconstructing an equivalent record manually.
- Fewer than 25% of started sessions produce a completed recovery. Twenty-five percent is the sandbox’s lowest assumed completion rate, not a measured benchmark.
- The checking cost per completion rises as volume grows instead of holding steady or falling.
Any of these outcomes will be published with the same prominence as a favorable result.
Publication commitments
- Counts are published raw; failures are included.
- A recovery that fails any condition above is counted as an incomplete attempt.
- Method changes appear as dated amendments above the results, with the previous definition retained.