Transparency

Escalations

Questions the society could not settle for itself. An escalation is raised when a decision needs authority, context or judgement the system does not have — a conflict between sources, a claim it cannot ground, a limit it must not cross on its own. They are published open, including the ones nobody has answered.

Evidence chain verified. 1424 records, head 85702a792688998a…. Every action is hash-chained to the one before it; a reader can recompute the chain from the raw log below.

381 open · 0 resolved · 381 total. An open escalation is not a defect in the log; it is the log working. A system that never escalates is either never uncertain or not reporting it.

Open (381)

idreasonquestionraised
esc-1critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-2passages_drifted2 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-3grounded_recovery_lowOnly 0.5385 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-4extraction_unstableThe two reader passes agreed on only 0.4118 of claims. Should extractions below this threshold be treated as unusable?
esc-5cross_paper_overlap1 claim(s) in 2606.19887 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-6passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-7checker_citation_not_found1 of 23 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-8grounded_recovery_lowOnly 0.6667 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-9extraction_unstableThe two reader passes agreed on only 0.6875 of claims. Should extractions below this threshold be treated as unusable?
esc-10prior_art_uncertainIs claim "Safety alignment suppresses engagement with sensitive knowledge, making it difficult for models to identify and reason about multimodal risk elements." already answered by prior work? Retrieval returned 9 candidates and the adjudicator could not decide.
esc-11checker_citation_not_found2 of 26 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-12grounded_recovery_lowOnly 0.6111 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-13extraction_unstableThe two reader passes agreed on only 0.6111 of claims. Should extractions below this threshold be treated as unusable?
esc-14cross_paper_overlap3 claim(s) in 2606.09711 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-15detail_not_in_cited_evidenceClaim 2606.09711/c1 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-16detail_not_in_cited_evidenceClaim 2606.09711/c2 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-17detail_not_in_cited_evidenceClaim 2606.09711/c3 names "Source B", "Source C", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-18detail_not_in_cited_evidenceClaim 2606.09711/c4 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-19detail_not_in_cited_evidenceClaim 2606.09711/c5 names "PRIME", "Source A", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-20detail_not_in_cited_evidenceClaim 2606.09711/c6 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-21detail_not_in_cited_evidenceClaim 2606.09711/c7 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-22detail_not_in_cited_evidenceClaim 2606.09711/c9 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-23detail_not_in_cited_evidenceClaim 2606.09711/c11 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-24detail_not_in_cited_evidenceClaim 2606.09711/c12 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-25detail_not_in_cited_evidenceClaim 2606.09711/c13 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-26detail_not_in_cited_evidenceClaim 2606.09711/c16 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-27detail_not_in_cited_evidenceClaim 2606.09711/c17 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-28detail_not_in_cited_evidenceClaim 2606.09711/c18 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-29detail_not_in_cited_evidenceClaim 2606.09711/c19 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-30detail_not_in_cited_evidenceClaim 2606.09711/c20 names "100", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-31detail_not_in_cited_evidenceClaim 2606.09711/c22 names "PRIME", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-32passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-33grounded_recovery_lowOnly 0.5 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-34extraction_unstableThe two reader passes agreed on only 0.5217 of claims. Should extractions below this threshold be treated as unusable?
esc-35cross_paper_overlap8 claim(s) in 2606.02630 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-36detail_not_in_cited_evidenceClaim 2606.02630/c3 names "four", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-37detail_not_in_cited_evidenceClaim 2606.02630/c14 names "six", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-38prior_art_uncertainIs claim "The paper characterizes four degradation trajectory signatures (Compliance Creep, Diminishing Returns, Pattern Recognition, Spike-and-Abandonment) that describe" already answered by prior work? Retrieval returned 18 candidates and the adjudicator could not decide.
esc-39prior_art_uncertainIs claim "Turn 2 is the critical vulnerability window for safety intervention in multi-turn medical conversations." already answered by prior work? Retrieval returned 23 candidates and the adjudicator could not decide.
esc-40passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-41grounded_recovery_lowOnly 0.6111 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-42extraction_unstableThe two reader passes agreed on only 0.5714 of claims. Should extractions below this threshold be treated as unusable?
esc-43cross_paper_overlap15 claim(s) in 2512.00349 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-44cross_paper_overlap8 claim(s) in 2606.17478 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-45detail_not_in_cited_evidenceClaim 2606.17478/c4 names "GPT-OSS-20B", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-46detail_not_in_cited_evidenceClaim 2606.17478/c11 names "LatentQA", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-47prior_art_uncertainIs claim "Threshold OR ensembles of monitor families reduce false negatives but raise realized Alpaca-control FPR." already answered by prior work? Retrieval returned 13 candidates and the adjudicator could not decide.
esc-48passages_drifted2 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-49grounded_recovery_lowOnly 0.4167 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-50extraction_unstableThe two reader passes agreed on only 0.3684 of claims. Should extractions below this threshold be treated as unusable?
esc-51cross_paper_overlap7 claim(s) in 2606.10747 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-52prior_art_uncertainIs claim "Instruction-induced misalignment produces salient behavioral cues: when the fine-tuned model organism is paired with a risky system prompt, pure observation alr" already answered by prior work? Retrieval returned 12 candidates and the adjudicator could not decide.
esc-53passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-54grounded_recovery_lowOnly 0.4444 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-55extraction_unstableThe two reader passes agreed on only 0.4167 of claims. Should extractions below this threshold be treated as unusable?
esc-56cross_paper_overlap14 claim(s) in 2509.02655 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-57grounded_recovery_lowOnly 0.625 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-58extraction_unstableThe two reader passes agreed on only 0.5263 of claims. Should extractions below this threshold be treated as unusable?
esc-59cross_paper_overlap12 claim(s) in 2606.08682 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-60detail_not_in_cited_evidenceClaim 2606.08682/c12 names "AS-induced EM", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-61detail_not_in_cited_evidenceClaim 2606.08682/c13 names "AS-induced EM", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-62prior_art_uncertainIs claim "Among the tested injection layer groups on Qwen3.5-27B, layers 22-25 give the highest EM rate, while injecting into layers 24-25 yields near-zero emergent misal" already answered by prior work? Retrieval returned 12 candidates and the adjudicator could not decide.
esc-63critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-64passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-65checker_citation_not_found1 of 23 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-66grounded_recovery_lowOnly 0.6154 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-67extraction_unstableThe two reader passes agreed on only 0.4706 of claims. Should extractions below this threshold be treated as unusable?
esc-68cross_paper_overlap17 claim(s) in 2606.28863 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-69detail_not_in_cited_evidenceClaim 2606.28863/c12 names "RLHF", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-70detail_not_in_cited_evidenceClaim 2606.28863/c19 names "three", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-71detail_not_in_cited_evidenceClaim 2606.28863/c22 names "four", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-72prior_art_uncertainIs claim "The paper claims the triadic test partitions cases into in-class and out-of-class: an honest safety filter and incidental distribution shift fall outside, while" already answered by prior work? Retrieval returned 12 candidates and the adjudicator could not decide.
esc-73prior_art_uncertainIs claim "The paper reports that among its thirty documented cases no two share the same (trigger, swap, origin) triple and that cases distribute across twenty-two of the" already answered by prior work? Retrieval returned 12 candidates and the adjudicator could not decide.
esc-74prior_art_uncertainIs claim "The paper reports that of its thirty documented cases, twelve are upward swaps, twelve are downward, and six are lateral (persona switch), and that the default " already answered by prior work? Retrieval returned 12 candidates and the adjudicator could not decide.
esc-75grounded_recovery_lowOnly 0.48 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-76extraction_unstableThe two reader passes agreed on only 0.48 of claims. Should extractions below this threshold be treated as unusable?
esc-77cross_paper_overlap27 claim(s) in 2605.24197 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-78detail_not_in_cited_evidenceClaim 2605.24197/c1 names "Multi-agent", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-79detail_not_in_cited_evidenceClaim 2605.24197/c12 names "Multi-agent", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-80detail_not_in_cited_evidenceClaim 2605.24197/c20 names "RewardBench", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-81detail_not_in_cited_evidenceClaim 2605.24197/c27 names "six", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-82prior_art_uncertainIs claim "Under epsilon-close priors and likelihoods and a sufficiently informative evidence lower bound, role posteriors remain delta-close; consequently, without distin" already answered by prior work? Retrieval returned 18 candidates and the adjudicator could not decide.
esc-83critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-84passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-85checker_citation_not_found20 of 137 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-86grounded_recovery_lowOnly 0.5926 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-87extraction_unstableThe two reader passes agreed on only 0.5714 of claims. Should extractions below this threshold be treated as unusable?
esc-88cross_paper_overlap10 claim(s) in 2603.00829 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-89detail_not_in_cited_evidenceClaim 2603.00829/c7 names "Kimi K2", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-90detail_not_in_cited_evidenceClaim 2603.00829/c12 names "LLMs", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-91prior_art_uncertainIs claim "Attempts to improve on grid-search-selected prompts via additional iterative refinement (human or automated) generally do not yield further gains and instead in" already answered by prior work? Retrieval returned 17 candidates and the adjudicator could not decide.
esc-92prior_art_uncertainIs claim "A grid search over 3 candidate models and 15 candidate prompts yields monitors with test-set partial AUROC of 0.853 (Gloom) and 0.866 (STRIDE)." already answered by prior work? Retrieval returned 11 candidates and the adjudicator could not decide.
esc-93prior_art_uncertainIs claim "Human-guided prompt refinement on STRIDE yields a statistically significant improvement over the best prompt-sweep prompt, an isolated exception to the general " already answered by prior work? Retrieval returned 9 candidates and the adjudicator could not decide.
esc-94checker_citation_not_found3 of 45 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-95grounded_recovery_lowOnly 0.3571 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-96extraction_unstableThe two reader passes agreed on only 0.2632 of claims. Should extractions below this threshold be treated as unusable?
esc-97cross_paper_overlap11 claim(s) in 2606.11409 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-98detail_not_in_cited_evidenceClaim 2606.11409/c5 names "HarmBench", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-99detail_not_in_cited_evidenceClaim 2606.11409/c7 names "PAIR", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-100detail_not_in_cited_evidenceClaim 2606.11409/c8 names "Qwen3-4B-SafeRL", "50", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-101prior_art_uncertainIs claim "Gradient-based GCG suffixes optimized on an open-weight surrogate (Qwen2.5-0.5B-Instruct) can transfer to a separate target model (Qwen3-8B), eliciting non-triv" already answered by prior work? Retrieval returned 16 candidates and the adjudicator could not decide.
esc-102passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-103checker_contradictedClaim "Safety-aligned RL on Qwen3-4B raises aggregate adversarial compute cost for JailBroken and PAIR while leaving some harm categories disproportionately exploitabl" was judged contradicted by a blinded checker reading only the passages cited for it. Decide whether it should remain in the graph.
esc-104checker_citation_not_found1 of 22 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-105grounded_recovery_lowOnly 0.4167 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-106extraction_unstableThe two reader passes agreed on only 0.3571 of claims. Should extractions below this threshold be treated as unusable?
esc-107cross_paper_overlap13 claim(s) in 2606.20626 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-108cross_paper_overlap15 claim(s) in 2501.14940 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-109detail_not_in_cited_evidenceClaim 2501.14940/c6 names "Contextual Integrity", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-110detail_not_in_cited_evidenceClaim 2501.14940/c15 names "CASE-Bench", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-111prior_art_uncertainIs claim "Each task (query-context pair) was annotated by 21 annotators, a number determined by statistical power analysis, and the dataset contains 47,000+ human annotat" already answered by prior work? Retrieval returned 16 candidates and the adjudicator could not decide.
esc-112prior_art_uncertainIs claim "The paper applies Contextual Integrity (CI) theory parameters to formalize context, describing this as the first instance of using CI theory to build a foundati" already answered by prior work? Retrieval returned 17 candidates and the adjudicator could not decide.
esc-113passage_fabricated1 passage(s) the reader cited do not appear in the source paper at all. Is the extractor inventing text, or is the PDF extraction mangling honest quotes?
esc-114passages_drifted2 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-115checker_citation_not_found1 of 25 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-116grounded_recovery_lowOnly 0.5714 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-117extraction_unstableThe two reader passes agreed on only 0.6471 of claims. Should extractions below this threshold be treated as unusable?
esc-118cross_paper_overlap16 claim(s) in 2606.00027 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-119prior_art_uncertainIs claim "The highest-scoring domains were Safety & Reliability and Medical Errors (each averaging around 0.96), while Bias, Fairness & Equity (0.95 ± 0.04 SD) and Clinic" already answered by prior work? Retrieval returned 12 candidates and the adjudicator could not decide.
esc-120prior_art_uncertainIs claim "Operationally complex categories including Liability, Accountability, and Medical Coding & Billing were the most challenging (domain means between 0.79 and 0.83" already answered by prior work? Retrieval returned 19 candidates and the adjudicator could not decide.
esc-121grounded_recovery_lowOnly 0.6667 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-122extraction_unstableThe two reader passes agreed on only 0.5455 of claims. Should extractions below this threshold be treated as unusable?
esc-123cross_paper_overlap18 claim(s) in 2606.04435 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-124detail_not_in_cited_evidenceClaim 2606.04435/c3 names "Confidence Inflation Cascade", "Context Poisoning", "Inference Cascade", "Poisoning Cascade", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-125detail_not_in_cited_evidenceClaim 2606.04435/c19 names "HITL-AP", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-126detail_not_in_cited_evidenceClaim 2606.04435/c21 names "Confidence Inflation Cascade", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-127passage_fabricated4 passage(s) the reader cited do not appear in the source paper at all. Is the extractor inventing text, or is the PDF extraction mangling honest quotes?
esc-128passages_drifted2 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-129checker_citation_not_found6 of 48 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-130grounded_recovery_lowOnly 0.375 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-131extraction_unstableThe two reader passes agreed on only 0.3913 of claims. Should extractions below this threshold be treated as unusable?
esc-132cross_paper_overlap12 claim(s) in 2606.05391 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-133detail_not_in_cited_evidenceClaim 2606.05391/c8 names "Co-planning", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-134detail_not_in_cited_evidenceClaim 2606.05391/c11 names "two", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-135passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-136grounded_recovery_lowOnly 0.4118 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-137extraction_unstableThe two reader passes agreed on only 0.3889 of claims. Should extractions below this threshold be treated as unusable?
esc-138cross_paper_overlap13 claim(s) in 2606.12918 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-139detail_not_in_cited_evidenceClaim 2606.12918/c2 names "Shapley-guided", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-140detail_not_in_cited_evidenceClaim 2606.12918/c4 names "2", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-141detail_not_in_cited_evidenceClaim 2606.12918/c6 names "Agent-level Shapley", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-142prior_art_uncertainIs claim "The paper designs a closed-loop, Shapley-guided autonomous red-teaming agent that selects a coalition of agents and jointly generates coordinated, role-aware ad" already answered by prior work? Retrieval returned 22 candidates and the adjudicator could not decide.
esc-143prior_art_uncertainIs claim "Prior red-teaming methods (TAMAS, GCA, AutoTransform, AiTM) yield near-zero attack success rates in most settings on hierarchical MAS, especially under limited " already answered by prior work? Retrieval returned 14 candidates and the adjudicator could not decide.
esc-144prior_art_uncertainIs claim "Agent-level Shapley value distributions are highly skewed and task-dependent: only a small subset of agents contributes significantly to attack success, and the" already answered by prior work? Retrieval returned 10 candidates and the adjudicator could not decide.
esc-145grounded_recovery_lowOnly 0.4615 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-146extraction_unstableThe two reader passes agreed on only 0.4286 of claims. Should extractions below this threshold be treated as unusable?
esc-147cross_paper_overlap14 claim(s) in 2606.05566 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-148detail_not_in_cited_evidenceClaim 2606.05566/c8 names "JBB-Behaviors", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-149prior_art_uncertainIs claim "The system operates with an average latency of approximately 50 ms on CPU, making it suitable for production deployment under cost and infrastructure constraint" already answered by prior work? Retrieval returned 13 candidates and the adjudicator could not decide.
esc-150prior_art_uncertainIs claim "The protectai-v2 model (184M parameters) achieves a perfect F1 of 1.000 on the awall-test benchmark but collapses to F1 = 0.000 on the unseen JBB-Behaviors pool" already answered by prior work? Retrieval returned 16 candidates and the adjudicator could not decide.
esc-151critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-152grounded_recovery_lowOnly 0.375 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-153extraction_unstableThe two reader passes agreed on only 0.2609 of claims. Should extractions below this threshold be treated as unusable?
esc-154cross_paper_overlap14 claim(s) in 2606.05233 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-155detail_not_in_cited_evidenceClaim 2606.05233/c1 names "158", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-156detail_not_in_cited_evidenceClaim 2606.05233/c4 names "Anthropic", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-157detail_not_in_cited_evidenceClaim 2606.05233/c10 names "17", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-158critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-159checker_citation_not_found3 of 32 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-160grounded_recovery_lowOnly 0.3125 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-161extraction_unstableThe two reader passes agreed on only 0.3125 of claims. Should extractions below this threshold be treated as unusable?
esc-162cross_paper_overlap16 claim(s) in 2606.03810 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-163detail_not_in_cited_evidenceClaim 2606.03810/c17 names "three", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-164prior_art_uncertainIs claim "The paper reports that the regularization methods ACT and BCT produce larger effects than label-generation methods, strongly suppressing reward hacking and emer" already answered by prior work? Retrieval returned 13 candidates and the adjudicator could not decide.
esc-165critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-166passages_drifted4 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-167grounded_recovery_lowOnly 0.5385 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-168extraction_unstableThe two reader passes agreed on only 0.45 of claims. Should extractions below this threshold be treated as unusable?
esc-169cross_paper_overlap17 claim(s) in 2606.01322 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-170detail_not_in_cited_evidenceClaim 2606.01322/c1 names "TukaBench", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-171detail_not_in_cited_evidenceClaim 2606.01322/c5 names "Afri-JBB-Culture", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-172detail_not_in_cited_evidenceClaim 2606.01322/c6 names "Code-switched", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-173detail_not_in_cited_evidenceClaim 2606.01322/c7 names "Boundary Point Jailbreaking", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-174detail_not_in_cited_evidenceClaim 2606.01322/c11 names "EFUSED", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-175detail_not_in_cited_evidenceClaim 2606.01322/c13 names "JAILBROKEN", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-176passages_drifted4 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-177grounded_recovery_lowOnly 0.5333 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-178extraction_unstableThe two reader passes agreed on only 0.4737 of claims. Should extractions below this threshold be treated as unusable?
esc-179cross_paper_overlap27 claim(s) in 2506.04018 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-180detail_not_in_cited_evidenceClaim 2506.04018/c3 names "two", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-181prior_art_uncertainIs claim "Two conditions are required for behavior to qualify as misaligned: acting contrary to the deployer's intended goals (rather than following malicious instruction" already answered by prior work? Retrieval returned 12 candidates and the adjudicator could not decide.
esc-182prior_art_uncertainIs claim "Most evaluations use the pre-built InspectAI basic agent, a simple ReAct loop with task-specific tools and a reflective prompt." already answered by prior work? Retrieval returned 20 candidates and the adjudicator could not decide.
esc-183grounded_recovery_lowOnly 0.6333 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-184extraction_unstableThe two reader passes agreed on only 0.6333 of claims. Should extractions below this threshold be treated as unusable?
esc-185cross_paper_overlap17 claim(s) in 2606.24081 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-186detail_not_in_cited_evidenceClaim 2606.24081/c4 names "eleven", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-187prior_art_uncertainIs claim "Methods that must be reconstructed primarily from paper text (PGJ, R2A, Low-Effort) show larger deviations, with PGJ at 7.2% error and R2A at 16.1% error, attri" already answered by prior work? Retrieval returned 16 candidates and the adjudicator could not decide.
esc-188critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-189passages_drifted2 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-190grounded_recovery_lowOnly 0.3529 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-191extraction_unstableThe two reader passes agreed on only 0.3158 of claims. Should extractions below this threshold be treated as unusable?
esc-192cross_paper_overlap13 claim(s) in 2604.23130 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-193detail_not_in_cited_evidenceClaim 2604.23130/c2 names "Single-token-driven", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-194detail_not_in_cited_evidenceClaim 2604.23130/c7 names "17", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-195prior_art_uncertainIs claim "Hierarchical-linkage steering is the most selective and least effective of the three strategies, because its cluster-size constraint (merged cluster at most 50 " already answered by prior work? Retrieval returned 9 candidates and the adjudicator could not decide.
esc-196prior_art_uncertainIs claim "The harm-responsible features are largely prompt-specific: 17.4% of steered responses on original adversarial prompts received a higher harmfulness score than t" already answered by prior work? Retrieval returned 16 candidates and the adjudicator could not decide.
esc-197prior_art_uncertainIs claim "Among responses that began as non-harmful content (default score 1), 3.70% were driven to maximal harm (score 5) and a further 1.10% to score 4, so 4.8% of non-" already answered by prior work? Retrieval returned 16 candidates and the adjudicator could not decide.
esc-198grounded_recovery_lowOnly 0.4615 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-199extraction_unstableThe two reader passes agreed on only 0.3333 of claims. Should extractions below this threshold be treated as unusable?
esc-200cross_paper_overlap18 claim(s) in 2606.24014 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-201detail_not_in_cited_evidenceClaim 2606.24014/c1 names "80", "50", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-202detail_not_in_cited_evidenceClaim 2606.24014/c14 names "twelve", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-203prior_art_uncertainIs claim "Across the evaluated OpenAI models, alignment evaluation scores show weak positive cross-model correlation (mean Spearman's rho = 0.107) and the first principal" already answered by prior work? Retrieval returned 20 candidates and the adjudicator could not decide.
esc-204passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-205checker_citation_not_found1 of 33 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-206grounded_recovery_lowOnly 0.2941 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-207extraction_unstableThe two reader passes agreed on only 0.25 of claims. Should extractions below this threshold be treated as unusable?
esc-208cross_paper_overlap15 claim(s) in 2601.19072 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-209detail_not_in_cited_evidenceClaim 2601.19072/c10 names "97", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-210checker_contradictedClaim "Tree of thought is consistently the top-performing strategy, direct assessment is second best, multi-step reasoning and few-shot achieve lower performance, and " was judged contradicted by a blinded checker reading only the passages cited for it. Decide whether it should remain in the graph.
esc-211grounded_recovery_lowOnly 0.4667 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-212extraction_unstableThe two reader passes agreed on only 0.4667 of claims. Should extractions below this threshold be treated as unusable?
esc-213cross_paper_overlap15 claim(s) in 2602.02557 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-214detail_not_in_cited_evidenceClaim 2602.02557/c15 names "Text-transferred", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-215prior_art_uncertainIs claim "Text attacks achieve the highest average StrongReject (SR) score across the evaluated omni-models, revealing a text-centric vulnerability." already answered by prior work? Retrieval returned 10 candidates and the adjudicator could not decide.
esc-216grounded_recovery_lowOnly 0.5 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-217extraction_unstableThe two reader passes agreed on only 0.5 of claims. Should extractions below this threshold be treated as unusable?
esc-218cross_paper_overlap18 claim(s) in 2606.28332 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-219detail_not_in_cited_evidenceClaim 2606.28332/c2 names "General-purpose", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-220detail_not_in_cited_evidenceClaim 2606.28332/c15 names "Safe Helpfulness", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-221passages_drifted4 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-222grounded_recovery_lowOnly 0.5714 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-223extraction_unstableThe two reader passes agreed on only 0.4545 of claims. Should extractions below this threshold be treated as unusable?
esc-224cross_paper_overlap22 claim(s) in 2606.07631 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-225detail_not_in_cited_evidenceClaim 2606.07631/c2 names "EM-relevant", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-226prior_art_uncertainIs claim "Alignment-relevant trait directions are necessary for low-FNR detection: the alignment feature set reaches 2.2% FNR under RF while semantic and random 7D contro" already answered by prior work? Retrieval returned 17 candidates and the adjudicator could not decide.
esc-227passage_fabricated1 passage(s) the reader cited do not appear in the source paper at all. Is the extractor inventing text, or is the PDF extraction mangling honest quotes?
esc-228passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-229checker_citation_not_found1 of 44 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-230grounded_recovery_lowOnly 0.2857 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-231extraction_unstableThe two reader passes agreed on only 0.3043 of claims. Should extractions below this threshold be treated as unusable?
esc-232cross_paper_overlap14 claim(s) in 2603.26846 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-233detail_not_in_cited_evidenceClaim 2603.26846/c2 names "two", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-234detail_not_in_cited_evidenceClaim 2603.26846/c8 names "Instruct", "Llama-3", "3.3", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-235detail_not_in_cited_evidenceClaim 2603.26846/c9 names "Gaussian", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-236detail_not_in_cited_evidenceClaim 2603.26846/c12 names "Lagrange", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-237prior_art_uncertainIs claim "Among the four stability metrics, semantic entropy (SE) maintains the most consistent separability across CoT and Response, whereas PPL, Pmax, and Cosine Sim ar" already answered by prior work? Retrieval returned 15 candidates and the adjudicator could not decide.
esc-238prior_art_uncertainIs claim "CoT Monitor induces obfuscated reward hacking, paradoxically worsening Actual Deception while collapsing CoT Faithfulness." already answered by prior work? Retrieval returned 15 candidates and the adjudicator could not decide.
esc-239prior_art_uncertainIs claim "SAR retains general model capability, performing within normal fluctuation ranges and avoiding alignment tax or capability collapse." already answered by prior work? Retrieval returned 11 candidates and the adjudicator could not decide.
esc-240critic_high_severity2 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-241passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-242checker_citation_not_found2 of 21 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-243grounded_recovery_lowOnly 0.5385 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-244extraction_unstableThe two reader passes agreed on only 0.5 of claims. Should extractions below this threshold be treated as unusable?
esc-245cross_paper_overlap13 claim(s) in 2505.14289 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-246prior_art_uncertainIs claim "Effective adversarial semantics form a dense, continuous 'semantic attack space' in the model's latent representation rather than isolated sparse points, which " already answered by prior work? Retrieval returned 15 candidates and the adjudicator could not decide.
esc-247prior_art_uncertainIs claim "An 'alignment paradox' exists: models with more extensive alignment training sometimes show increased vulnerability to EVA's attacks, because alignment training" already answered by prior work? Retrieval returned 13 candidates and the adjudicator could not decide.
esc-248prior_art_uncertainIs claim "Successful adversarial payloads are not uniformly distributed over persuasion dimensions but concentrate on two attractors—trust-aligned and urgency-aligned sem" already answered by prior work? Retrieval returned 13 candidates and the adjudicator could not decide.
esc-249passages_drifted2 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-250checker_citation_not_found1 of 24 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-251grounded_recovery_lowOnly 0.7273 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-252extraction_unstableThe two reader passes agreed on only 0.6923 of claims. Should extractions below this threshold be treated as unusable?
esc-253cross_paper_overlap10 claim(s) in 2606.15396 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-254detail_not_in_cited_evidenceClaim 2606.15396/c4 names "three", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-255prior_art_uncertainIs claim "MDPO dynamically adjusts the KL penalty coefficient based on the policy model's real-time responsiveness to sample difficulty, using normalized reward gaps, out" already answered by prior work? Retrieval returned 15 candidates and the adjudicator could not decide.
esc-256critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-257passages_drifted2 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-258grounded_recovery_lowOnly 0.375 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-259extraction_unstableThe two reader passes agreed on only 0.2667 of claims. Should extractions below this threshold be treated as unusable?
esc-260cross_paper_overlap15 claim(s) in 2606.00033 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-261grounded_recovery_lowOnly 0.5882 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-262extraction_unstableThe two reader passes agreed on only 0.4167 of claims. Should extractions below this threshold be treated as unusable?
esc-263cross_paper_overlap14 claim(s) in 1805.03090 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-264detail_not_in_cited_evidenceClaim 1805.03090/c5 names "POMDPs", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-265grounded_recovery_lowOnly 0.2667 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-266extraction_unstableThe two reader passes agreed on only 0.2353 of claims. Should extractions below this threshold be treated as unusable?
esc-267cross_paper_overlap19 claim(s) in 2604.26360 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-268detail_not_in_cited_evidenceClaim 2604.26360/c13 names "Hopper-v4", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-269prior_art_uncertainIs claim "The reciprocal reliability filter is derived from risk-sensitive mean-variance utility and replaces an unbounded linear penalty that can become negative under h" already answered by prior work? Retrieval returned 15 candidates and the adjudicator could not decide.
esc-270prior_art_uncertainIs claim "The reciprocal reliability filter satisfies positivity, monotonicity, boundedness, identity at zero uncertainty, and Lipschitz continuity." already answered by prior work? Retrieval returned 15 candidates and the adjudicator could not decide.
esc-271prior_art_uncertainIs claim "The magnitude of reward misspecification is assumed to be bounded by a monotonically increasing function of epistemic and aleatoric uncertainty." already answered by prior work? Retrieval returned 13 candidates and the adjudicator could not decide.
esc-272passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-273checker_citation_not_found1 of 37 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-274grounded_recovery_lowOnly 0.4211 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-275extraction_unstableThe two reader passes agreed on only 0.4 of claims. Should extractions below this threshold be treated as unusable?
esc-276cross_paper_overlap14 claim(s) in 2606.07706 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-277critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-278grounded_recovery_lowOnly 0.7857 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-279extraction_unstableThe two reader passes agreed on only 0.7333 of claims. Should extractions below this threshold be treated as unusable?
esc-280cross_paper_overlap22 claim(s) in 2606.07612 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-281detail_not_in_cited_evidenceClaim 2606.07612/c16 names "Fine-tuning", "Instruct", "Llama-3", "3.1", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-282detail_not_in_cited_evidenceClaim 2606.07612/c19 names "3.15", "5.18", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-283passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-284checker_citation_not_found3 of 56 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-285grounded_recovery_lowOnly 0.2381 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-286extraction_unstableThe two reader passes agreed on only 0.2273 of claims. Should extractions below this threshold be treated as unusable?
esc-287cross_paper_overlap17 claim(s) in 2502.20914 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-288detail_not_in_cited_evidenceClaim 2502.20914/c13 names "Gaussian", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-289prior_art_uncertainIs claim "An exhaustive circuit-first pass on the illustrative example network (k = 3, n = 1, loss cutoff 10^-3) yielded 59 circuits and 114,230 interpretations." already answered by prior work? Retrieval returned 17 candidates and the adjudicator could not decide.
esc-290grounded_recovery_lowOnly 0.4737 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-291extraction_unstableThe two reader passes agreed on only 0.4737 of claims. Should extractions below this threshold be treated as unusable?
esc-292cross_paper_overlap12 claim(s) in 2604.24668 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-293detail_not_in_cited_evidenceClaim 2604.24668/c5 names "50", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-294detail_not_in_cited_evidenceClaim 2604.24668/c9 names "two", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-295passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-296grounded_recovery_lowOnly 0.4545 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-297extraction_unstableThe two reader passes agreed on only 0.375 of claims. Should extractions below this threshold be treated as unusable?
esc-298cross_paper_overlap17 claim(s) in 2606.18988 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-299detail_not_in_cited_evidenceClaim 2606.18988/c4 names "VAC-GRPO", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-300detail_not_in_cited_evidenceClaim 2606.18988/c15 names "four", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-301prior_art_uncertainIs claim "The authors construct Deception-10K, described as the first fine-grained audio-visual Chain-of-Thought dataset, comprising 10,000 video-reasoning pairs (~50 hou" already answered by prior work? Retrieval returned 18 candidates and the adjudicator could not decide.
esc-302prior_art_uncertainIs claim "The paper proposes Visual-Audio Consistency Group Relative Policy Optimization (VAC-GRPO) with a progressive training strategy that stratifies data into four di" already answered by prior work? Retrieval returned 17 candidates and the adjudicator could not decide.
esc-303prior_art_uncertainIs claim "Most existing baseline multimodal large language models perform around the random-guess baseline of 50% on deception detection despite identical prompts." already answered by prior work? Retrieval returned 13 candidates and the adjudicator could not decide.
esc-304critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-305passage_fabricated1 passage(s) the reader cited do not appear in the source paper at all. Is the extractor inventing text, or is the PDF extraction mangling honest quotes?
esc-306passages_drifted3 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-307grounded_recovery_lowOnly 0.5385 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-308extraction_unstableThe two reader passes agreed on only 0.4706 of claims. Should extractions below this threshold be treated as unusable?
esc-309cross_paper_overlap14 claim(s) in 2606.08451 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-310detail_not_in_cited_evidenceClaim 2606.08451/c8 names "six", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-311detail_not_in_cited_evidenceClaim 2606.08451/c13 names "38", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-312prior_art_uncertainIs claim "In the most severe cases, models agree with harmful prompts over 70% of the time in zero-shot languages, defaulting to explicit agreement with safety-critical p" already answered by prior work? Retrieval returned 20 candidates and the adjudicator could not decide.
esc-313critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-314passages_drifted2 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-315grounded_recovery_lowOnly 0.5833 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-316extraction_unstableThe two reader passes agreed on only 0.5 of claims. Should extractions below this threshold be treated as unusable?
esc-317cross_paper_overlap16 claim(s) in 2606.10106 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-318detail_not_in_cited_evidenceClaim 2606.10106/c8 names "four", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-319detail_not_in_cited_evidenceClaim 2606.10106/c9 names "four", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-320checker_citation_not_found4 of 34 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-321grounded_recovery_lowOnly 0.625 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-322extraction_unstableThe two reader passes agreed on only 0.5 of claims. Should extractions below this threshold be treated as unusable?
esc-323cross_paper_overlap19 claim(s) in 2606.07532 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-324detail_not_in_cited_evidenceClaim 2606.07532/c2 names "Justice", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-325detail_not_in_cited_evidenceClaim 2606.07532/c6 names "Experiment", "200", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-326detail_not_in_cited_evidenceClaim 2606.07532/c10 names "ChatEval", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-327prior_art_uncertainIs claim "A single-model ablation isolating identity stripping (Experiment 3 / pilot14) produces a directional accuracy gain of 6.0 percentage points over unstripped cont" already answered by prior work? Retrieval returned 13 candidates and the adjudicator could not decide.
esc-328checker_citation_not_found2 of 59 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-329grounded_recovery_lowOnly 0.6 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-330extraction_unstableThe two reader passes agreed on only 0.48 of claims. Should extractions below this threshold be treated as unusable?
esc-331cross_paper_overlap18 claim(s) in 2606.20814 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-332detail_not_in_cited_evidenceClaim 2606.20814/c2 names "Instruct", "Qwen2", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-333prior_art_uncertainIs claim "The relationship between evaluation sample size and the Max Score Difference roughly follows a power law." already answered by prior work? Retrieval returned 14 candidates and the adjudicator could not decide.
esc-334prior_art_uncertainIs claim "The raw (non-JSON, non-template) format of the Initial EM questions is almost always the most misaligned format." already answered by prior work? Retrieval returned 14 candidates and the adjudicator could not decide.
esc-335prior_art_uncertainIs claim "Even with variations in learning schedules, in-domain loss has a dominant effect on the level of misalignment in the paper's experiment setting." already answered by prior work? Retrieval returned 19 candidates and the adjudicator could not decide.
esc-336prior_art_uncertainIs claim "Across the training-dynamics experiments summarized in Table 3, training loss still guides the level of misalignment." already answered by prior work? Retrieval returned 22 candidates and the adjudicator could not decide.
esc-337passages_drifted2 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-338grounded_recovery_lowOnly 0.5625 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-339extraction_unstableThe two reader passes agreed on only 0.5556 of claims. Should extractions below this threshold be treated as unusable?
esc-340cross_paper_overlap14 claim(s) in 2606.26793 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-341detail_not_in_cited_evidenceClaim 2606.26793/c2 names "Prior Sampling", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-342critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-343checker_citation_not_found1 of 20 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-344grounded_recovery_lowOnly 0.625 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-345extraction_unstableThe two reader passes agreed on only 0.625 of claims. Should extractions below this threshold be treated as unusable?
esc-346cross_paper_overlap8 claim(s) in 2602.18008 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-347detail_not_in_cited_evidenceClaim 2602.18008/c2 names "Existing LLM-based", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-348prior_art_uncertainIs claim "Models generated by NIMMGen can be used for counterfactual intervention simulation: increasing simulated social distancing strength produces systematic reductio" already answered by prior work? Retrieval returned 21 candidates and the adjudicator could not decide.
esc-349checker_citation_not_found3 of 22 quote(s) produced by the checker could not be located in the material it was shown. Its verdicts rest on text it may not have read correctly.
esc-350grounded_recovery_lowOnly 0.75 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-351extraction_unstableThe two reader passes agreed on only 0.5 of claims. Should extractions below this threshold be treated as unusable?
esc-352cross_paper_overlap21 claim(s) in 2606.27188 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-353detail_not_in_cited_evidenceClaim 2606.27188/c1 names "Agentic Business Process", "Business Process Management", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-354detail_not_in_cited_evidenceClaim 2606.27188/c5 names "Model Context Protocol", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-355detail_not_in_cited_evidenceClaim 2606.27188/c16 names "two", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-356detail_not_in_cited_evidenceClaim 2606.27188/c18 names "CUGA FLO", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-357grounded_recovery_lowOnly 0.3333 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-358extraction_unstableThe two reader passes agreed on only 0.3333 of claims. Should extractions below this threshold be treated as unusable?
esc-359cross_paper_overlap27 claim(s) in 2605.11047 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-360prior_art_uncertainIs claim "Qwen3.5-Plus, DeepSeek-v4-Flash, and DeepSeek-v4-Pro show consistently high AGS across the six risk categories, indicating the generated traps transfer beyond t" already answered by prior work? Retrieval returned 11 candidates and the adjudicator could not decide.
esc-361passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-362grounded_recovery_lowOnly 0.4615 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-363extraction_unstableThe two reader passes agreed on only 0.4444 of claims. Should extractions below this threshold be treated as unusable?
esc-364cross_paper_overlap16 claim(s) in 2510.00845 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-365prior_art_uncertainIs claim "Bootstrap resampling of the input dataset yields the lowest structural consistency and highest variability of discovered circuits (Jaccard µ = 0.561, CV = 0.335" already answered by prior work? Retrieval returned 16 candidates and the adjudicator could not decide.
esc-366prior_art_uncertainIs claim "Circuits discovered under bootstrap resampling also have the highest average circuit error (0.440), meaning they are structurally different and less faithful to" already answered by prior work? Retrieval returned 12 candidates and the adjudicator could not decide.
esc-367prior_art_uncertainIs claim "Shifting the meta-distribution (meta-dataset or prompt paraphrasing) yields more stable circuits than bootstrap resampling, with higher Jaccard indices (0.790 a" already answered by prior work? Retrieval returned 15 candidates and the adjudicator could not decide.
esc-368prior_art_uncertainIs claim "Circuit discovery methods do not scale trivially: stability degrades for larger models, with gpt2-small yielding relatively clustered results while Llama-3.2 (1" already answered by prior work? Retrieval returned 9 candidates and the adjudicator could not decide.
esc-369passages_drifted1 passage(s) are neither verbatim nor invented: they open with real source text and then diverge into paraphrase. Should the extractor be required to quote contiguously, or should near-quotes be accepted with a marker?
esc-370grounded_recovery_lowOnly 0.5333 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-371extraction_unstableThe two reader passes agreed on only 0.5 of claims. Should extractions below this threshold be treated as unusable?
esc-372cross_paper_overlap26 claim(s) in 2606.21399 closely restate a claim the society already holds from another paper. Is this an independent confirmation, a duplicate that should be merged, or a contradiction that should be recorded?
esc-373detail_not_in_cited_evidenceClaim 2606.21399/c2 names "two", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-374detail_not_in_cited_evidenceClaim 2606.21399/c6 names "two", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-375detail_not_in_cited_evidenceClaim 2606.21399/c10 names "0.423", "0.394", "0.436", "0.417", "four", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-376detail_not_in_cited_evidenceClaim 2606.21399/c15 names "ScienceWorld", "23", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-377detail_not_in_cited_evidenceClaim 2606.21399/c17 names "Prompt-only", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-378detail_not_in_cited_evidenceClaim 2606.21399/c20 names "On ALFWorld", which is real paper vocabulary but appears in none of the passages cited for this claim, nor in the window shown to the checker. Is this misattribution or partial citation?
esc-379critic_high_severity1 high-severity objection(s) stand against this extraction. Do any invalidate a claim that should not remain in the graph?
esc-380grounded_recovery_lowOnly 0.2692 of source-grounded claims were recovered by the second reader pass. Is this coverage diversity, or extractor unreliability? The distinction determines whether the 0.50 raw overlap is bad news.
esc-381extraction_unstableThe two reader passes agreed on only 0.2692 of claims. Should extractions below this threshold be treated as unusable?

Machine-readable: /society/escalations.json.