The evidence boundary in automated research
An automated research system can move from literature retrieval to a proposed idea, experiments, and a manuscript. Yet the evidence supplied to it is only a subset of prior research. RINI studies evidence-boundary overclaiming: describing a contribution as more novel than verified prior work supports because decisive evidence was absent from the system’s context.
The question is not whether an answer sounds original. It is whether a specific novelty claim remains defensible after the relevant evidence is made available.
Separate the intervention from natural retrieval
The controlled study compares the same research direction under matched evidence conditions: neutral filler (W), same-topic but non-covering evidence (H), and decisive prior work (E). Holding the generator, prompt, seed, and evidence-slot budget fixed helps isolate the role of the evidence. The primary H–E comparison distinguishes decisive prior work from merely seeing more material on the same topic.
A separate natural-retrieval cohort investigates whether retrieval misses the relevant document or the decisive passage inside it. Evidence insertion into genuine misses addresses a different question from the controlled study; the two cohorts are not pooled into one effect estimate.
Measure claims, not an impression of originality
The manuscript defines Unsupported Novelty Burden (UNB) from a complete inventory of atomic claims. It measures the positive gap between asserted novelty and the novelty supportable by a verified, frozen prior-evidence set, normalized for output length. It is not a universal originality score or a test of whether a paper appeared in model training.
Retrieve, certify, and repair
RINI Dual searches through original, rewritten, challenger, and differentiator queries, then verifies candidate evidence in full text under a fixed retrieval budget. A Minimal Reversal Certificate (MRC) identifies the smallest verified evidence set sufficient to narrow an unsupported claim while retaining defensible differences.
The repair step targets the affected claim rather than replacing the entire research contribution with a generic disclaimer. Acceptance checks preserve research intent, specificity, attribution, and supported extensions; failed checks send the case for review. The final assessment remains separate from the repair gate.
Source basis and status
This brief follows the current working manuscript, RINI: When Missing Evidence Masquerades as Originality, especially its abstract, introduction, and method. It describes the study protocol and proposed retrieval-and-repair system. Formal experiment and annotation exports are still pending; no completed effect estimates or validated repair gains are claimed here. Public paper and code links will be added when released.

Loading comments…