effectcheck 0.7.17

Welch t-test effect sizes (changes verdicts on Welch rows with a stated N far above df + 2 only)

effectcheck 0.7.16

Parser (changes rows on markdown input and on rho-subscript eta-squared only)

Worker: submit, then poll

Frontend

QA

effectcheck 0.7.15

Worker only: no file under R/ changed. The worker’s analysis code moved into worker/jobs.R unchanged, and responses from the new background worker are byte-identical to the in-process path (tested on all six analysis routes). What changes is what happens when the server is asked to do two things at once.

Frontend follow-up, added after the worker release (no file under R/ or worker/ changed; the package version stays 0.7.15 because nothing the R package or worker returns changed).

Concurrent requests are refused with a retryable 503 instead of queueing into a 502 (d-8fa65b). The R worker is single-threaded. Measured on the live free-tier worker (2026-08-09): three concurrent /api/v1/process-text calls finished at 98s / 120s / 120s, two came back as HTTP 502 from the proxy, and /health was unanswerable for ~120s. Every analysis route now runs its work in a background R process (future + promises, new worker/jobs.R), so the HTTP process stays free. While an analysis runs, a further request to any analysis route is refused at once with 503, Retry-After: 15 and {"busy": true, "reason": "server_busy", ...}. The refusal has no error field on purpose: a client that treats an unknown error code as fatal (SciMeto does) would otherwise give up on a document that was never looked at. /health now answers during an analysis and reports an admission block (mode, workers, in_flight, refused_since_start, …).

effectcheck 0.7.14

Ten open defects, re-tested on 0.7.13 first; the ones that still reproduced are fixed here. Every new test was watched failing on 0.7.13. Several published values change – listed in full so a consumer can tell a fix from a regression.

A confidence-interval verdict of INCONSISTENT now requires an interval comparable with the paper’s. ci_check_status = "INCONSISTENT" says the paper’s interval is wrong; 0.7.9 withdrew its status escalation after it accused 3 correct papers out of 3, but the column kept publishing INCONSISTENT on the same false premises. It is now UNVERIFIABLE, with the new column ci_unverifiable_reason saying why, when the row itself shows the intervals are not comparable:

MATCH and PLAUSIBLE are never altered (except generalized eta-squared). A genuinely wrong interval of each shape stays INCONSISTENT – each rule has that control in test-v0714-ci-verdict-premise.R. ci_match is NA (not FALSE) on these rows. (d-cf56d6, d-4efdc1)

A Welch test’s effect size is no longer “verified” against an N solved from it. With group sizes unstated a Welch df only bounds N from below, so the Welch branch back-solved N = 4t2/d2 from the reported d and then recomputed d from that N – reproducing it by construction. Measured at 0.7.13: a deliberately wrong d = 0.70 on Welch's t(223) = 8.11 (true about 0.99) returned PASS at N = 537. Such a row now checks its p-value only, says the effect size is not verified, and publishes N_source = "effect_backsolved" (it said global_text or not_found). When the back-solve falls below the Welch floor the N is clamped to df + 2, the check there is real, and the label is now df_inferred. Stated group sizes now give N = n1 + n2 (N_source = "group_sizes") instead of being ignored in favour of the back-solve: Welch's t(222.87) = 8.11, d = 0.99 with n1 = 131, n2 = 135 now verifies d at N = 266 as PASS. An integer Welch df no longer produces the note “Non-integer df (223.00)”. (d-48266c; the filed symptom, N = 225 and a false WARN, was already gone at 0.7.13.) Two further Welch fixes from the /ship review (2026-09-30): a Welch test compares independent groups, so paired-design variants (dz, dav, drm) are no longer offered on a Welch row – a wrong d = 0.70 with n1 = 131, n2 = 135 had passed against drm = 0.685 (0.7.13 too); and a tiny, rounded Welch d that the Welch floor N = df + 2 reproduces within rounding is graded at that N instead of a scraped study total. Measured on the 27-paper seam (docpluck 2.4.147): 6 rows move, all in 10.1016/j.jesp.2020.104052 and 10.5334/irsp.571 – the five jesp rows now grade at N = 171-203 (df_inferred) instead of the document total N = 827 (3 NOTE -> PASS), irsp.571 stays WARN against the independent-groups d. A row whose effect is explicitly labelled dz, dav or drm keeps its paired variants even when its df is fractional or its clause mentions Welch: the tier-2 cross-model check showed a correct dz = 0.44 on a within-subject Satterthwaite contrast t(45.3) = 3.00 falling from PASS to NOTE under the Welch gate.

Generalized eta-squared is no longer published as partial eta-squared. The generalized_eta2 column and all_variants computed it with the partial formula whenever the design was unclear – on collabra.126266 it equalled partial_eta2 on all 15 F rows beside a row message saying it cannot be computed. It depends on which factors are measured or within-subjects (Olejnik & Algina, 2003; Bakeman, 2005), which F and df do not carry, so it is now always NA. A row reporting it no longer carries two false notes (“unusual for F-test”, “symbol unclear – may be OCR”), checks its p-value (it was SKIP, “nothing checked”), and its interval is not graded against partial-eta-squared intervals (collabra.126266: 8 INCONSISTENT, 5 PLAUSIBLE, 1 MATCH, all against the wrong estimand). The 9 bare abstract restatements the gold counts carry no test statistic; effectcheck extracts no bare effect size of any family, so that is a scope question, logged (d-368d63). (d-b1c193)

A sentence reporting both the average direct effect and the ACME yields both. Chan & Feldman (2025, doi 10.1080/02699931.2024.2434156) print “The average direct effect was 0.15, 95% CI [-0.13 to 0.45], p = .3, whereas the … indirect effect (ACME) was 0.67, 95% CI [0.47-0.89], p < .001”. 0.7.13 emitted one row: the direct effect, labelled indirect_effect; the ACME was absent. Now two rows, direct_effect and indirect_effect. A sentence mixing one CI-form and one Sobel-form effect still yields one row (pre-existing, logged d-564e1b). (d-4fe38e)

A statistic quoted twice is one result. 10.24072/pci.rr.100726 is a review letter quoting t(868) = -3.01, p = .006 twice to show comma placement; both copies were scored. v0.6.18 withdrew three dedup rules because two distinct results can share every printed number, and named the missing signal: quoted material. Rows now merge only when the printed signature is identical, they are adjacent, and every copy is inside quotation marks; the kept row says so. The v0.6.18 guard-rail tests (distinct results must never merge) still pass. The N = 870 provenance and the 2.2x p-value discrepancy on that paper were already fixed at 0.7.13. (d-43ed61)

An undecidable post-hoc contrast names the contrast reading. collabra.90203’s t(998) = 0.097, d = 0.01 sits under a Bonferroni post-hoc announcement with an omnibus F(2, 998); at a d that small, N = 1000 and the contrast N of about 667 both fit, so the incumbent N = 1000 stands – and the row now says the contrast reading exists. The filed defect (the decisive contrasts binding N = 1000, a false WARN and two false CI flags) was already fixed at 0.7.13. (d-43c69d)

Already fixed at 0.7.13, closed on re-test: collabra.57785’s ambiguous-design t(742) row publishes N = 743 (gold 743) and names both readings (d-bd5e80); cog_emo Table 9’s three false CI-mismatch flags are gone – the rows keep N = 794 under the 2026-09-04 ruling to report, not silently correct, an ambiguous sample, and the one remaining INCONSISTENT is the paper’s own dropped minus sign (d-4ea375).

The API returns every column on every row (worker). JSON endpoints serialized rows with jsonlite’s data-frame default, which drops a key whose value is NA; a row carried 62 keys where check_text() returns 138. Every row now carries every column, NA as null. No other response field changed (measured by diffing the full process-text response). A consumer that detects schema drift will log each newly visible key once. (d-c7c216)

Harness: the gold lock is promoted to article-finder’s 2026-09-25 page-checked correction of 10.1016/j.jesp.2021.104154 (81 -> 81 results); run_validation.R refused on origin/main without it.

effectcheck 0.7.13

Rows from a table no caption claimed are no longer checked. docpluck 2.4.145 (released and live in production on 2026-09-27, while 0.7.12 was serving) starts including rows from grids that no table caption claims – docpluck measured about half of those as page furniture such as flow-diagram boxes and author blocks – and marks them caption_status = "uncaptioned_candidate". Measured 2026-09-27 on the 27 papers of the end-to-end seam set, effectcheck 0.7.12 against local docpluck 2.4.144 and the 2.4.145 candidate: 85 result rows gained, 0 lost, PASS 377 -> 377 and OK 268 -> 268 (the new rows verified nothing), and two new false WARNs on 10.1016/j.jesp.2022.104372, where a questionnaire-item grid was read as t-tests on 2 and 6 degrees of freedom, each with an INCONSISTENT CI. Those rows are now dropped before any other table-row step, so an uncaptioned grid cannot displace a captioned table that repeats it. On that paper the output is again exactly what docpluck 2.4.144 produces. The docpluck boundary golden is refreshed at the v2.4.145 tag (deb69ae): the t-test keys t, d, df that 2.4.144 dropped are back, the fixture yields the 5 rows the page prints (the 2.4.141 golden recorded 8, two of them invented), and caption_status is the one new field. The contract gate is green again; the owner waiver recorded in 0.7.12 is closed. The boundary contract now also ASSERTS which docpluck build answered (fleet thread T-0009): the golden’s recorded docpluck_identity must equal the live service’s, and the live version_crosscheck must read agreed. Before this the identity block was write-only, so a docpluck release that renamed no field passed the contract silently; now it fails until the capture is reviewed and the golden refreshed (worker/tests/test-docpluck-contract-identity.R). Nothing changes on docpluck 2.4.144, which does not send these rows. A real statistical table printed without a caption is therefore not checked – as it is not today. The count set aside is recorded as n_table_rows_uncaptioned_dropped in the result’s settings (a Sonnet review pointed out that n_table_rows alone counted rows that were never looked at).

The docpluck version-comparison harness (tests/harness/diff_endtoend_versions.py) no longer crashes on a Windows console when a row label holds a character outside cp1252.

Documentation rewritten and kept honest by a gate. The README now covers every exported function with its real signature, every check_text() argument, the six statuses and the result attributes, and docs/output-columns.md describes all 140 result columns (the README previously listed 18). Stale claims corrected: the README cited version 0.5.7; it documented filter_by_source(x, sources) and filter_by_delta(x, min, max), whose arguments are files, pattern, min_delta and max_delta; it described the effectsize package as the primary computation engine, but its code path in compute.R has an empty body and is never reached (interval methods come from MBESS and analytic formulas); and the vignette still showed the defunct file functions and a three-status table. Four check_text() arguments – tol_p, sign_sensitive, ci_method_phi, ci_method_V – are accepted and recorded in the settings but read by no computation; this was measured (changing each leaves the full result identical, while changing tol_effect or ci_affects_status does not) and is now stated in their help and the README rather than implied away. tools/check_docs_coverage.R derives the public surface from the code and fails when any export, argument, value, column, setting or option is missing from the docs, or when the README quickstart does not run against a fresh install; it is pinned two-sided by tests/testthat/test-docs-coverage-gate.R. Added CITATION.cff and CONTRIBUTING.md. No behaviour change.

effectcheck 0.7.12

A table row the extractor typed as an impossible test no longer publishes that test. docpluck 2.4.143 and 2.4.144 (the version production serves) attach the rows printed below a ruled table to that table, so on a page of stacked tables the next table’s rows are read under the first table’s header. On the boundary-contract fixture a t-test row Sleep vs Control 2.41 98 .018 0.15 arrived as F = 98, df1 = 0.018, df2 = 0.15, and 0.7.11 published F(0.018, 0.15) = 98 – on a SKIP row, which does not suppress values. A structured table row is now refused when its typed test cannot exist: an F with numerator df below 1, a non-positive denominator df or a negative F; a t with a non-positive df. These are the statistics’ domains, not plausibility limits – no Greenhouse-Geisser, Huynh-Feldt, Welch or mixed-model df crosses them. (A t or F-denominator df between 0 and 1 is deliberately admitted: a mixed-model Satterthwaite df can legitimately print one.) The row is kept, so a reader can see that a table result could not be read, but every number on it is withheld; it is a NOTE with extraction_suspect = TRUE and two new columns, df_guard_rejected and df_guard_reason, saying why. Kept rather than dropped because a dropped row makes a paper with an unreadable table look fully checked.

One table delivered twice is checked once. The same upstream bug returned Table 1’s three F rows a second time inside the misattached table, labelled “Table 2”, so each result was checked and counted twice, once under the wrong table. The rule is whole-table containment: a later table’s rows are dropped only when EVERY result row of an earlier table on the same page (at least two distinct rows, each with at least two numbers one of which is a t, F or r) reappears in it field for field – row label, group, every typed value including p_op. Only the repeated rows are dropped; the first occurrence, in delivery order, is kept. Two genuine tables that agree on some rows, a repeat on another page, rows without a table id or a page – all kept, the conservative direction.

Prose gets the same domain, flagged rather than withheld. An F with df1 below 1 or df2 at or below 0 in body text is now extraction_suspect with an “IMPOSSIBLE DF” reason (a df at or below 0 was already flagged); its values stay visible because the reader has the source sentence beside them. The existing df1 message printed its value with %.0f, so a df1 of 0.018 read “df1 = 0”; it now prints the value as read.

Prevalence, measured 2026-09-25 on the end-to-end corpus through a clean local docpluck 2.4.144: 27 papers, 26 extracted with the table step ok (10.5334/irsp.945 extracts zero characters, a known separate defect). Across 2,578 flattened table rows – 54 typed F rows, 179 typed t rows – there is no fractional or non-positive F numerator df, no non-positive denominator df, no negative F and no t df below 1; the final code refuses 0 rows and drops 0 duplicates, and the 990 result rows check_text() builds from those papers carry df_guard_rejected = FALSE throughout. The same code on the same service refuses 2 rows and drops 3 duplicates on the contract fixture, so the zeros are real zeros, not a guard that cannot fire. The defect shape needs stacked tables on one page with a ruled table above; none of the corpus papers delivered it, which is also why no corpus-level gate caught it.

Measured against a clean local docpluck 2.4.144 before the fix: the boundary contract exits 1 with flattened_row_field_keys: REMOVED/RENAMED upstream: d, df, t and 9 rows where the golden has 8. The golden is deliberately not refreshed: the upstream fix (docpluck commit 28e0b8f) is on docpluck’s main branch but in no tagged release yet, and the contract says to refresh only once the t/df/d keys return. Note the golden itself is not ground truth: it records 8 flattened rows where the page prints 5 results, i.e. it was captured with Table 1 already delivered twice (at 2.4.141). A correct upstream fix will therefore move the count 8 -> 5, and that change must not be read as a regression.

Released with the boundary contract RED, by explicit owner waiver (2026-09-25). /escicheck-qa reported the contract BROKEN against docpluck 2.4.144 (d, df, t keys missing; 9 rows vs 8). That red describes what production docpluck sends whether or not this release ships, and this release is the mitigation for it; the only way to turn it green from this side – refreshing the golden to 2.4.144 – would record the broken rows as correct. The waiver covers exactly that one red and nothing else. docpluck 2.4.145 (announced, not yet tagged) fixes the rows and also adds caption_status to every flattened row, including rows from uncaptioned grids (about half page furniture); ESCImate must decide how to treat caption_status == "uncaptioned_candidate" rows before production moves to 2.4.145. The first draft was reviewed by two other models (Sonnet, then Sol), and every defect they found was reproduced by a test that failed first: single matches on different pages pooling into a “block of two” (Sonnet); a t-df floor of 1 that would have hidden a legitimate mixed-model result, and a “two matching rows” dedup rule that could merge two genuine tables and missed a partial third copy (Sol) – which is why the rule is now full containment. New regression test test-v0712-table-row-domain-and-cross-table-dup.R.

effectcheck 0.7.11

A row from which no effect size could be recomputed no longer claims a cross-family fallback that never ran. When a results table prints F, p, η²p and a CI but no degrees of freedom (10.1525/collabra.90203 Table 8), nothing can be recomputed from the F, so the set of computed variants is empty. The row nevertheless said “No same-type variants available for ‘etap2’ - using all computed variants [category: cross-family]” – describing a fallback to variants that did not exist, under the category API.md defines as “the matcher cross-falls to the closest computed variant in a different family”. It now reads “No effect-size variants could be computed from this row for ‘etap2’ (e.g. its degrees of freedom are not reported), so the reported effect size was not compared to anything [category: not-computed]”. The same correction applies when the effect-size type itself was not stated.

Text only, deliberately. ambiguity_level stays "highly_ambiguous", so design_ambiguous, confidence, uncertainty_level and status are unchanged on every row; setting the level to "clear" instead would have raised confidence by 6 points on rows where nothing was verified. A genuine cross-family fallback (e.g. F(2, 30) = 5.00, d = 0.60) keeps its [category: cross-family] tag. [category: not-computed] is a new, third tag; consumers splitting on the two existing tags see it as neither.

Found by the escicheck-iterate canary audit of 10.1525/collabra.90203 on 2026-09-23, confirmed in triage by three models (Opus, Sonnet 5, Fable 5). The same audit’s seven reported findings were all triaged to other owners or to documented behaviour – none was an effectcheck defect: the missing partial-eta-squared labels are drawn as vector shapes in the PDF and never reach the text (a docpluck extraction limit); the absent “Target article” rows are the original study’s statistics, excluded on purpose since 0.6.6; and design_ambiguous = FALSE on prose F rows is the documented contract (it describes matching ambiguity, which design_inferred does not). New regression test test-v0711-no-variants-not-cross-family.R.

effectcheck 0.7.10

R CMD check --as-cran regained a 0-warning build: one em-dash literal in a check.R uncertainty message (introduced by 0.7.9’s CI-escalation reword) tripped “checking code files for non-ASCII characters” – the check flags non-ASCII in R code (string literals, identifiers) but tolerates it in comments, so the rest of that same message and every other non-ASCII comment in the file were never the issue. Replaced with ASCII --; no message text or behaviour changed beyond the glyph. New regression test test-ascii-source-discipline.R runs tools::showNonASCIIfile() over effectcheck/R/ so this class cannot recur silently.

This release exists as a SEPARATE version rather than as an amendment to 0.7.9, and that is the point. 0.7.9 was already serving in production, and /health reports packageVersion("effectcheck") – so shipping a changed R/check.R under the same number would have made two different builds both answer 0.7.9, leaving deploy-drift-check.sh and /ship’s own Phase 5.0 “did the deploy land?” gate structurally unable to distinguish them. That is the v0.7.3 failure shape, where a rolled-back release answered status:healthy for 15 hours. A version string is the only identity those gates have; changing the artifact without changing the string disarms them silently.

No parsing, computation, or verdict logic changed in this release beyond the ASCII fix above. The AI-gold regeneration of three papers (10.1371__journal.pmed.1004323, 10.1098__rsos.250908, 10.1016__j.joep.2020.102349) that also landed this cycle is test-CORPUS work, not a code change – it corrected ground truth that had been transcribed in the parser’s own notation rather than the paper’s printed form (see TODO.md 2026-09-08/09 and tests/harness/gold.lock.json history[]), and moved the replay harness’s accuracy readout from 28.8% to 29.7% CORRECTLY VERDICTED with zero behavioural change on this side of the seam.

effectcheck 0.7.9

A confidence-interval check that fired only on correct papers was built, measured, and withdrawn inside one release. It escalated a row to WARN when the reported interval matched none of the intervals this package computes, on the premise that the candidate universe had been exhausted so the paper must be wrong. Measured over 704 rows of real published text (the 49-paper corpus in article-finder custody, 0.7.8 against the escalating build, 219 duplicate keys dropped identically from both arms), it fired three times and every one was a correct paper – two of which had been PASS. In each case the author’s interval had been graded against an interval for a different quantity: a paired dz interval for a between-groups Welch d, and a Spearman interval for a paper whose own sentence reads “A Pearson’s correlation was computed”. Each interval was recomputed independently of this package from the paper’s own numbers, recovering n by inverting the t test where unstated, and all three match to three decimals ([-0.00, 0.34] vs [-0.0018, 0.3418]; [-0.060, 0.539] vs [-0.0597, 0.5390]; [-0.033, 0.545] vs [-0.0324, 0.5456]).

Five independent ways the premise fails were found by three model providers – cross-family scale, an assumed confidence level, an assumed equal-N group split, a one-sided interval, and a CI whose referent is not the reported effect. Two providers found the split case independently of each other. One exemption was written for the first; the next four arrived within the hour. The ways “we computed a comparable interval” can be false are open-ended, which is an inverted default rather than a list of missing special cases. A check that fires on correct input is worse than no check: an author who follows a flag and finds nothing behind it learns to ignore the next one, and the next one may be real.

Nothing was lost by withdrawing it, because the severity was already published. ci_check_status grades MATCH / PLAUSIBLE / INCONSISTENT / UNVERIFIABLE / MISSING on every row and always has, and ci_method_match names the method the interval was actually compared against – the field that makes an INCONSISTENT interpretable, and the one that reveals the Pearson-graded-against-Spearman case. The complaint that motivated the escalation was that a downstream consumer maps status only and ignores both. API.md now says to read them, and records the measurement so this is not re-proposed. status behaviour for a mismatched interval is identical to 0.7.8.

The one genuine catch is kept, in a form that cannot misfire. A reported interval lying outside its effect’s mathematical range – a correlation outside [-1, 1], an eta-squared above 1 – now joins the impossible_value family beside the reversed-interval check, reusing the same bounds table so the two cannot drift. It compares the paper against a bound rather than against anything computed here, so it is immune to all five premise failures. The control that defines its scope: a correlation interval of [0.90, 0.99], badly wrong for r = .34 but possible, correctly stays PASS.

API.md documented five row statuses and omitted SKIP, which the code emits. A consumer building a status map from the published documentation wrote exactly the incomplete map behind a downstream “insufficient data” defect; a second consumer confirmed the same gap. All six are now documented, with what SKIP means (check.R: an extraction-only row with nothing checked and nothing worth surfacing) and an explicit note that OK verifies the p-value, never the effect size – all six OK write sites are in the p-value branch, and PASS is the status that structurally requires a matched value.

.effectcheck_version() no longer falls back to a hardcoded "0.2.0" (deferred from the previous release): it returns explicit not-installed / unreadable markers instead, so a lookup failure can never be mistaken for a real version. No published artifact was ever affected – the worker’s version lookup is an unguarded top-level call at startup, so any process that served a result had already proved the lookup succeeded.

Also in this release: the project’s own cleanup skill read as permission to delete without asking (its delete recipe sat ~30 lines above the rule reserving deletion to the user, and it named a gate string printed by a different command than the scans that produce the hit list), and the merged fix-queue work from the sibling branch – a robust-family crash, an estimate-outside-CI check, CI symmetry on correlations, correlation N-provenance reaching the output, and a chi-square back-solve that could not fail.

effectcheck 0.7.8

A build that reported success shipped a different image, and a confidence interval changed under a method label that did not move. Both were measured on 2026-09-02 while verifying – for the first time – that the 2026-08-09 build pin does what its comment claims. It does: a local rebuild three weeks later reproduced production’s engine set exactly (R 4.6.1, effectsize 1.0.3, MBESS 5.0.1, stringi 1.8.9, stringr 1.6.0, ICU 74.2) and the 30-document corpus came back byte-identical, 505 rows / 2,343,217 normalized characters. What was NOT deterministic was the build around it.

Two cold-cache builds of the identical Dockerfile produced 158 and 140 packages. remotes::install_local(dependencies = TRUE) planned 49; in the second build 18 failed to download, each emitting only Warning: download of package 'x' failed. R continued, the RUN step exited 0, and the image shipped without statcheck, testthat, shiny, DT, ggplot2 and 13 others. Nothing caught it. /health was byte-identical to the healthy image, because statcheck was not among the packages it names; and the test suite could not run, because testthat was one of the casualties. The user-visible consequence, measured: POST /api/v1/compare-text returned one row instead of two – statcheck’s entire second opinion silently absent, HTTP 200, no error. The core statistical path was unaffected (that image also produced 30/30 byte-identical corpus output), and production was verified healthy, so nothing published was wrong. The exposure was the next rebuild.

The build now fails rather than shipping a quietly different image: an explicit required-package assertion after install_local, proven three ways – it passes on a healthy image (21 required of 158 installed), fails when a package is removed, and fails on the actually-degraded image naming statcheck and testthat.

ci_d_ind() reported noncentral_t for bounds it did not compute that way. ci_d_ind_noncentral_t() falls through to ci_d_ind_approx() – a large-sample approximation – on two paths: |ncp| beyond R’s noncentral-t accuracy limit (~37.62), and MBESS not installed. Both were labelled noncentral_t. Measured in two fresh R processes on d = 0.67, n1 = n2 = 25: [0.0964533619, 1.2369531589] with MBESS, [0.0996671786, 1.2403328214] without – the interval moved and the label did not. MBESS is a Suggests, installed through the step that was dropping packages silently, so this was reachable from a build that reported success. Bounds now carry a ci_engine attribute and the label is derived from it. uniroot_nct still reports noncentral_t because it is a noncentral-t inversion; only the genuinely approximate engine reports differently, so every row that was correct before is byte-identical after – verified on the full 30-document corpus, 0 changes.

Also: statcheck is now reported in /health’s engine_versions, derived rather than hand-listed by a test that scans the computation sources for Suggests called via ::; CRAN_SNAPSHOT moved from ARG to a committed literal, closing a one-flag --build-arg hole in the line whose whole purpose is to be un-overridable; and an ICU assertion pins the one part of the deliberately-unpinned apt layer that can move a published number (measured inert across both builds, but ICU drives stringi’s Unicode normalization, which is the exact mechanism behind the 49,091 vs 50,101 character incident). Suite 1236 test_that blocks across 145 files.

effectcheck 0.7.7

A retry that skipped its own second chance, a test that guarded a branch production never runs, and a comparison endpoint that quietly published NA. No parser, effect-size or verdict logic changed in this release: every fix is in the worker’s transport and provenance plumbing, plus the release gates that were supposed to catch these and did not.

The occasion was the v0.7.6 consumer-readiness handoff, which listed three changes that had shipped asserted only by source grep – a 429 + Retry-After retry whose retry never executed in any test, the removal of sections=true proved by grepping our own source rather than the query string, and extraction_provenance with zero worker tests on the one expression that carries it to production. Closing those gaps found real defects behind two of the three.

Fixed

Tests

Worker suite 46 passed / 2 skipped -> 145 passed / 0 skipped. Every new assertion was watched failing against a real mutation of the code it guards, including reverting the retry to its previous sequential-if form (three tests red) and removing the NA cap guard (R’s missing value where TRUE/FALSE needed).

The end-to-end provenance test now drives a real multipart request through plumber’s own router (pr$call()), so the CORS filter, body parser, route match, handler and serializer all run. The previous version built req$FILES by hand and passed – and scanning plumber 1.3.3, webutils 1.2.2 and httpuv 1.6.17 symbol by symbol finds "FILES" zero times: nothing in the serving stack sets it. That test was green on a branch production never takes.

Package suite 1232 test_that blocks / 3745 assertions / 0 failures, across 144 files; no package logic changed in this release. The one added block is the boundary-contract canary described under Release gates below.

Release gates

New scripts/verify-upstream-and-consumers.mjs, wired into all four project skills, with checks for: an unread upstream docpluck outbox (a FAIL, not a note – three went unread for eight days and one broke production for seven); corpus provenance and completeness; the symbol-contract canary; the consumer pin table in the new docs/CONSUMERS.md; and a release notification for the version being shipped.

Three of its own checks were false greens on first review and are fixed: SKIP exited 0, the canary hash check matched the word “SHAPE” in a comment, and the pin table was validated against itself rather than against each consumer’s source.

All testing now runs against a LOCAL docpluck service. The harnesses refuse a non-loopback URL unless --allow-remote is passed, because .Renviron points DOCPLUCK_URL at the metered hosted endpoint and any script that merely read the environment billed production silently.

The docpluck BOUNDARY contract. The symbol-contract snapshot pins what docpluck declares; this pins what it actually sends. Two committed synthetic fixtures are pushed across the wire to a local docpluck and the returned field set is diffed against a committed golden (inst/docpluck-contract/boundary_contract_golden.json, the one new file this release ships to CRAN users). The pair is the point: on 2026-08-14 the declared table and the wire agreed with each other and disagreed with effectcheck – eta2p became eta2_p, every partial eta-squared was dropped, and a real statistical inconsistency published as a clean PASS for seven days. Nothing in the declaration was wrong, so a declared-contract check structurally could not see it.

Exit 0 PASS / 1 FAIL / 2 COULD-NOT-VERIFY; a run without a local docpluck never reports green. Verified two-sided against a synthetic capture: renaming a probe’s VALUE while its KEY is unchanged – the 2026-08-14 defect exactly – is caught and both strings printed, and a removed key is caught and labelled REMOVED/RENAMED.

The golden is currently HELD at docpluck 2.4.137 / normalization 1.9.58. The local 2.4.138 build adds four keys and removes none, and docpluck confirms that re-goldening against a dirty build is forbidden because dirty is not a reproducible identity. It refreshes when 2.4.138 is tagged.

.gitattributes was added in the same change and is load-bearing: both fixtures are hashed by the golden, and with core.autocrlf=true a checkout rewrote their bytes, so the recorded hashes matched on the machine that captured them and on no fresh clone. Measured with git checkout-index before the fixtures were ever committed.

effectcheck 0.7.6

Two live defects that turned a real verdict green, and the six our extractor filed against us that nobody had read. docpluck had sent THREE outboxes since v0.7.5 and only the newest reached this repository; the two older ones carried the larger changes.

A partial eta-squared stopped parsing on ~2026-08-14 and the row went GREEN. docpluck’s symbol contract v2.0 (shipped v2.4.130) _-joins every subscript run, so eta2p became eta2_p and omega2p became omega2_p. Neither was in effectcheck’s alternation. Measured on the released 0.7.5: eta2_p = .11 published effect_reported = NA at status OK, where the identical eta2p = .11 scored ERROR — the effect size was dropped AND the verdict flipped from ERROR to a clean pass. omega2_p did the same from WARN. Scope was measured rather than assumed: those two tokens are the ONLY ones that broke; eta2, eta2G, omega2, epsilon2, R2, f2, chi2 and both p < 10^-8 and p < 10-8 were unaffected.

Restored page boundaries reverted the v0.7.4 chunk fix. docpluck v2.4.136 stopped destroying form feeds (its page-number strip used \s, which matches U+000C, so it ate the page break beside the number it deleted). effectcheck had no form-feed handling at all. Measured through check_text(): two results separated by \n\f\n collapsed into ONE row carrying the first result’s d = 0.33 against the second’s t = 7.47 — a pairing that appears in no paper, and precisely the defect v0.7.4 shipped to remove. normalize_text() also fabricated dz = 3 from a page number, the v0.6.20 bridging class.

The fix DELETES the form feed early, before the whitespace collapse and both number strips. Deletion rather than translation was measured, not reasoned: \f -> \n turns the corpus-majority \n\f into \n\n and manufactures a chunk split that never existed, and one cross-model reviewer recommended exactly that. A form feed was never invisible here — PCRE \s matches it, so the sentence splitter already broke on one and the =-joiner already bridged across one. What it was not was a PARAGRAPH.

Impossible values are now refused rather than computed with. pat_SE has always accepted a leading sign and nothing checked it, so a negative standard error flowed into the t = b/SE synthesis and inverted the statistic — and verify_t_from_b_SE, the one check that exists to validate that synthesis, absolutises BOTH sides and is structurally incapable of noticing. A negative SE is now rejected before the division, with extraction_suspect and a named reason; the row is still emitted, because suppressing it would restore the silent-loss class v0.6.20 removed. A reversed reported interval (ciL > ciU) joins the impossible-value family, and the b-scale CI message no longer claims a field is “absent” when it was present and refused. The docpluck flattened-rows path now computes p_valid / p_out_of_range from the value instead of hardcoding them.

The occasion is docpluck’s rule W0g, which infers a missing minus from arithmetic. Reproduced on THIS repository’s own corpus by diffing docpluck 2.4.136 with the rule disabled: frontiers_music_mood_2024 had SE = 0.199 rewritten to -0.199 and p = 0.069 to -0.069; efendic_2022_affect had [0.22, 0.75] rewritten to [-0.22, 0.75] and [0.04, 0.25] to [0.04, -0.25]. These guards are NOT about W0g and do not expire with it — a standard error is non-negative and an interval is ordered, whatever produced them.

Upstream provenance is now visible. check_text() gains extraction_provenance; the worker forwards docpluck’s normalization block, which it had received all along and never passed on (steps_changed had zero occurrences in this package before this release). Two new columns, upstream_sign_rewrites and upstream_normalization_version, report whether the extractor REWROTE VALUES in this document. Keyed on docpluck’s metric key, never on a rule name, so it degrades to a no-op when they delete the rule. DOCUMENT-level and deliberately NOT wired to extraction_suspect: that flag gates effect-size decimal REWRITING and two ERROR-path downgrades, so raising it on every row because one span was rewritten would demote unrelated genuine inconsistencies.

The six defects docpluck filed against 0.7.5 on 2026-08-13, every one reproduced at HEAD before being touched and verified red on the released 0.7.5: AF [6, 7] was rewritten into F-test notation because the bracket rule had no letter lookbehind (an F-test fabricated out of prose); the outline stripper deleted a published value (90.6 Third-plus generation ...) it could not tell from a heading; v2.1.451,52. fused into v2.1451.52.; 9999999,1 converted as a decimal; a space-separated affiliation run became Frank 1.2; and ~25%6,28 became %6.28.

The seventh is the one worth repeating. docpluck read scored 0,87 failing to convert as a false negative in our coded|dummy|scored vocabulary, and was about to adopt that vocabulary because of it. The vocabulary was never the cause: each element of the protected run is a single \d, so the pattern matched the PREFIX 0,8. Measurement stopped a sibling project adopting a mechanism for a reason that was not true.

Two of these fixes needed a second pass, both caught by testing rather than by reading the pattern: the first outline-stripper guard still deleted values, and the first affiliation guard used \s*, and therefore also protected Median 0,45, SD 0,12, suppressing a real European decimal.

Normalization spec 1.4.0 -> 1.5.0, 72 -> 80 conformance cases, and its status changes. docpluck DELETED its entire EU->US separator machinery (A3, A3a, A3c, A3d, A2, W0n, and the whole document-level locale feature) in v2.4.129-v2.4.130 — verified against the live library, which now delivers d = 0,80, U = 12,345 and N = 185,178 verbatim. effectcheck is therefore the ONLY implementation, the spec is no longer a two-implementation contract, and REQUEST_TO_DOCPLUCK_normalization_spec.md is moot rather than pending. One consequence is the opposite of what was feared: v0.7.5’s locale inference now works BETTER on docpluck output, because docpluck no longer touches the tokens it votes on.

Cost and resilience. sections=true is no longer requested — it has been sent on every extraction since v0.6.4 and consumed by nothing. A 429 is now retried once, bounded, honouring Retry-After (which had zero occurrences anywhere in this repository). /health reports the docpluck release behind the most recent result, and it reports the one that can be trusted. metadata.docpluck_version resolves from an environment variable inside docpluck’s own frontend and fails two ways: UNSET gives the literal string "unknown" (measured against a local instance that was running 2.4.136), and STALE gives a real-looking wrong number — production reported "2.4.101" while returning normalization.version = "1.9.57", which is 2.4.136’s normalization version. The first draft of this guard excluded only "unknown" and therefore published the stale 2.4.101 as fact. It was caught by the post-deploy probe, not by any local gate — no fixture can contain a value only production knows. normalization.version is computed from the library and cannot drift from it, so it is authoritative; the self-reported string survives as a labelled fallback (self-reported-<x>) for the case where no normalization report exists at all.

Corpus evidence. Whole-corpus diff over 37 real-article texts, all v0.7.6 parser changes against unchanged input: 0 rows gained, 0 lost, 2 verdict changes — both the new reversed-CI guard firing on genuinely reversed published intervals, and the 2026-08-13 fixes adding no further change at all. One was confirmed against the rasterized source page: Frontiers in Psychology 10.3389/fpsyg.2024.1303262 p6 prints (r = 0.195, p = 0.240, CI 95%[0.485, 0.132]). That row previously passed.

Known limitation, disclosed rather than discovered later: the regression corpus was found to be raw pdftotext, not docpluck output — 38 of its 48 files came from the pre-v0.4.0 read_any_text() path, so every whole-corpus diff quoted for v0.7.4 and v0.7.5 was measured on text production does not consume. Re-extraction through the production HTTP path is underway and tracked separately; this release does not claim it is finished.

effectcheck 0.7.5

A backlog item that would have shipped a feature doing nothing, a p-value scraped off a figure legend, and a locale signal three rules computed and none read. v0.7.5 closes the 2026-08-09 handoff. Two of its defects were found by checking the premise before writing the code, and one by verifying a fix against the article text rather than against the test that had just gone green.

Issue C – the resample count B, and the feature that would have been inert

The Monte Carlo floor check shipped in v0.6.22 (p >= 1/(B+1), Phipson & Smyth 2010) had never fired on any paper. Measured across the 48-file validation corpus before any change: ten rows carried a resampling_method and resampling_B was non-NA on zero of them. resampling_p_below_floor was FALSE everywhere – not because every paper passed, but because the check had no B to test against. A check that cannot fire is indistinguishable from one that passes.

The handoff attributed this to B being declared once in Methods, and prescribed a Methods prescan. That is half of it. The other half only appears by reading the paper: PNAS 10.1073/pnas.2404157121 does not write B = 10,000 anywhere. It writes “For permutation tests, ten thousand random shuffles of labels … were sampled” – the count is SPELLED OUT, and a qualifier sits between it and its noun. A prescan for the digit form would have added a helper, a provenance value and a test suite, and still bound nothing on the one paper it was written for. Of the four resample-count declarations in the corpus, the previous clause-level scan could read exactly one.

So the fix is a document-level Methods prescan (.doc_resampling_b()) plus a bounded number-word reader, bootstrap\w* on the qualifier, and a generic noun admitted when a resampling word is elsewhere in the same sentence. New column resampling_B_source (own_clause / methods_prescan), and every message that quotes B now says which. The PNAS paper’s six resampling rows carry B = 10000 and the floor check is live.

The first draft of this feature shipped the exact defect the feature exists to prevent, and the corpus scan caught it. brjpsych_1.txt contains (b=0.81, z=2.80, p=0.005, OR=2.25, ...); the case-insensitive B = <num> form matched b=0.81 and then stripped the . – correct for a thousands separator, catastrophic for a decimal – giving B = 81. A floor of 1/82 = 0.0122 would have declared every p below .0122 in that paper unattainable. Three guards now: case-sensitive uppercase B, an integral-shape requirement so 0.81 and 2.25 cannot pass, and a resampling word required in the same sentence.

A whole-document scan was measured against the Methods-scoped one and refused on the evidence: it gains one correct bind (collabra.126266, whose declaration genuinely sits in Results) and two wrong ones, including B = 60 scraped off a grading scale (Grade A+=80% or above, A=70-79%, B=60-69%). The positional scope pays for itself. collabra.126266 stays uncovered, deliberately, with the reason recorded.

Issue D – the permutation p was being discarded

A clause can report two p-values of different provenance: t(2037) = -3.26, P = 0.001, P-permutation = 0.002. pat_p binds the parametric 0.001 – it cannot see the hyphenated form at all – and 0.002 was thrown away, so a reader of the output could not tell that a permutation p had been reported.

New sibling column p_reported_secondary (plus p_secondary_symbol), never a second row: a new row would change nrow() for every consumer and silently shift every downstream index, and downstream’s field registry is frozen at v0.4.0, so it already tolerates unknown columns and cannot tolerate unknown rows. Scoped to the GLUED qualifier only, which is the whole safety argument – a glued qualifier is provably invisible to pat_p, so what this captures is provably not what pat_p bound. A spaced qualifier may already BE the primary, and capturing it could publish the same number twice under two provenances.

The floor / missing-B / Monte-Carlo-SE caveats now retarget to whichever p came out of the resampling distribution, and name it (“Reported permutation p (0.002) is below the minimum attainable …”). Previously the entire block was skipped on exactly these rows.

A p-value scraped off a figure legend – found by checking the source

Verifying Issue D against the article text turned up a defect in the same paper. The PNAS figure caption is merged into the body by the extractor, ends ...***P < 0.001., and continues in LOWERCASE, so the chunk splitter – which needs a capital or a digit – cannot separate them. pat_p took the FIRST p in the merged chunk and the row published t(2037) = -2.19, p_reported = 0.1, status WARN, where the paper prints P = 0.029 (and 2*pt(-2.19, 2037) = 0.02864, so the correct value is consistent and the row is OK). A threshold from an asterisk key, attached to a real published statistic, with no flag.

pat_p now refuses a p preceded by a legend marker (*, ~, +, #, either spacing). Fixed at the p-binding rather than at the chunk boundary deliberately: a boundary rule would have to guess where the caption ends, while this states something simply true, and it cannot separate a statistic from its own values (invariant 6). 80 occurrences across 12 of the 48 corpus papers.

Whole-corpus diff: 0 rows gained, 0 lost, THREE changed – all three this same defect in three different papers, each new value checked against the article:

paper was is legend
pnas_cognitive_memory_2024 p = 0.1 (WARN) p = 0.029 (OK) ~P < 0.1
frontiers_retrocue_2024 p = 0.05 p = 0.271 *p<0.050
scireports_exercise_2025 p = 0.01 p = 0.021 # P<0.01

Locale conflict was computed by one function and read by none

infer_numeric_locale() returns decisive / none / conflict and sets decimal_mark = NA for two of them. Three separate rules gated on identical(decimal_mark, ","), which collapses conflict into none, so a document that actively contradicts itself was normalized as if it were decisively US. Reproduced: a document mixing p = .035, d = 0.80 with Welch's correction gave t(2,758) = 3,21, d = 0,45 stripped the comma and published df = 2758, N = 2760 and a computed d = 0.122 against a reported 0.45 – a false WARN carrying a fabricated effect size on a correctly reported result. Under the decisive-European branch the identical string yields “cannot verify” with no computed value; conflict now reaches that same outcome.

The third of those rules is one the v0.7.3 cross-model audit had already fixed for the European case – by writing a fourth hand-maintained copy of the same test, which is precisely why the third state got past all of them. They now share .locale_comma_unresolved(). Corpus diff for this change alone: 0 of 764 rows change, because 0 of 48 papers infer conflict. That emptiness is the evidence the change is safe, not a reason it was skippable – a computed signal nothing reads is a trap, because it reads as handled at every site that mentions it.

Shared normalization spec 1.3.0 -> 1.4.0, 70 -> 72 conformance cases (new rule L3). SPEC.md’s own version stamp had been stale at 1.0.0 for three minor versions; corrected.

New test type mean_diff_ci – estimate + CI + p, no test statistic

ieee_access_alt reports three results as Mean difference of 457.66 articles, p-value = 7.171e-11, confidence interval (320.98, 594.35) and produced zero rows: every pattern anchors on a test statistic or a standardized effect size and this clause has neither. The triple is nonetheless mutually checkable, and md_hl already establishes that a row may carry CI-symmetry and p-CI-consistency checks while claiming no effect size. Scope confirmed with the user before implementing.

Three checks, and no effect size is ever claimed (a mean difference “in articles” has no standardizer in the clause; inventing one would be the cross-scale defect ci_referent exists to prevent):

  1. the estimate must be the midpoint of its own interval – a normal-approximation CI is symmetric by construction, unlike a Hodges-Lehmann one;
  2. the interval implies SE = (U - L)/(2z), and SE plus the estimate imply a p-value, compared on a ratio – these p-values are of order 1e-11, where any absolute tolerance waves everything through (the v0.6.18 lesson);
  3. p and the interval must agree on the significance decision.

The correctly reported paper is NOT flagged: implied 5.291e-11 against a reported 7.171e-11, a factor of 1.36.

Two supporting fixes, both general rather than scoped to the new type:

Corpus diff: 0 rows lost, 0 changed, 3 gained, all in this previously zero-row paper.

Build determinism (P0 from the 2026-08-07 portfolio survey)

Dockerfile:1 was the only literal :latest base image in the portfolio, and effectsize and MBESS – which compute the effect sizes and confidence intervals – installed from a moving repository. The same commit could publish a different CI on two different days with nothing to attribute the change to.

The Docker build is not verified locally (no Docker daemon available in this session). Every input was verified against its primary source – the registry digest, the p3m snapshot URL, and the base image’s own setup_R.sh – but the first real proof is the deploy, and the version gate below is what will surface a failure instead of hiding it.

Deploy and CI

Filed to docpluck, not fixed here

DP-14 (the 2026-08-09 extraction-tool defect log): bmj_1’s supplementary table arrives column-major – one cell per line, no row grouping, and a header split mid-word (Odds Ratio or Coefficien / t for Treatmen t Group). Reassembling it means guessing which estimate pairs with which interval, and a wrong guess there is not a missing row but a fabricated one – the v0.7.4 two-column defect and DP-11 both. ESCImate correctly renders nothing rather than something plausible and wrong.

Cross-model review — 12 findings, 8 confirmed and fixed, 4 refuted

Codex (gpt-5.5) and Claude Sonnet reviewed the diff independently. Every finding was REPRODUCED locally before being acted on, and every refutation was reproduced before being dismissed. Two of the confirmed defects were in code this release added to PREVENT that exact defect class, which is the argument for the review in one line.

Confirmed and fixed (8). Seven concern .resample_count_in() / the Methods prescan, where a wrong B produces a Monte-Carlo floor of 1/(B+1) and therefore a FALSE ACCUSATION against a correctly reported p-value:

# input that bound a wrong B was
1 "Bootstrap analyses were not used; vitamin B = 60 mg was administered." B = 60 — the sentence-level resampling gate is satisfied by a NEGATED mention
2 "Grade A+=80%, A=70-79%, B=60-69%, C=50-59%" B = 60 — the UNSPACED grading scale; the first fix only refused the spaced B = 60 mg form
3 "Each condition was tested with 60 replicates." B = 60 — replicates is ordinary wet-lab vocabulary and its branch was ungated
4 "We enrolled 240 iterations of the survey; a permutation test followed later." B = 240 — the loose noun accepted a resampling word anywhere in the sentence, in any order
5 a count in an Appendix after a Methods heading adopted — Appendix was not a closing heading
6 a count after 3.1 Sample Characteristics adopted — and no numbering rule can fix it, because normalize_text() strips section numbers before the prescan runs (verified, after a first fix that assumed otherwise and did nothing)

Every count form now requires a resampling word in its own sentence — the unambiguous nouns satisfy that themselves, so the gate costs nothing — the loose noun additionally requires the qualifier to PRECEDE the number and sit near it, B = <n> must have nothing else claiming the number, and the Methods region closes on any short standalone heading-like line rather than on a vocabulary that can always be incomplete.

  1. mean_diff_ci false-fired on ordinary rounding. The CI-symmetry check used a fraction of the interval width, and a correctly reported narrow interval at 2 dp exceeds it for no reason but rounding: 0.06, 95% CI [0.02, 0.09] gives a relative asymmetry of 0.143 against a 0.02 threshold, on numbers that are exactly consistent before rounding. The tolerance is now ROUNDING, not a fraction: each value is rounded to its own last place, so the midpoint can legitimately differ from the estimate by about one unit there. A genuinely mis-stated estimate is still caught.

  2. The bullet chunk rule could sever a chi-square from its own odds ratio. "chi2(1) = 12.74, p = .013. - N = 211, OR = 0.99, 95% CI [0.77, 1.27]." degraded from mcnemar_or carrying OR = 0.99 [0.77, 1.27] to a bare chisq with both NA — invariant 6, and exactly how collabra.37122 lost an odds ratio in v0.7.4. A hyphen is also a dash, a minus and a range, so the marker alone is not evidence; the boundary now additionally requires that what FOLLOWS the marker is not a statistic assignment (N =, OR =).

    Removing - from the marker set was tried first and silently reverted ieee_access_alt from 3 rows to 1: normalize_text() maps the U+FFFD the extractor delivers to a HYPHEN before chunking, so - was the load-bearing member all along and this file’s own comment crediting U+FFFD was wrong. Now asserted per-marker so it cannot recur unnoticed.

Refuted, and recorded rather than “fixed” (4).

One finding is accepted as a deliberate trade-off, not fixed: the legend guard also refuses a per-result marginal-significance marker (+p = .08). The corpus has 80 legend occurrences and ZERO alnum-preceded ones, and the two failure directions are not symmetric — refusing loses a p (the row reports NA and checks nothing), while admitting publishes a legend threshold AS a reported p-value. Losing a value is recoverable; fabricating one is not.

Blast radius of all eight fixes: 0 rows gained, 0 lost, 0 changed across the 48-paper corpus. They are pure tightening on shapes that do not occur in it — which is the point: they were found by construction, not by the corpus, and the corpus is what proves they cost nothing.

Verification

effectcheck 0.7.4

A column boundary was being read as part of a sentence, so an effect size from one column was graded against a test statistic from another. The chunk splitter broke text on [.!?] followed by whitespace and an uppercase letter. A two-column PDF whose columns the extractor merges resumes with whatever the next column happens to start with — often a bare number — and the lookahead required a capital, so the two columns stayed one chunk.

spps.txt location 216 shipped

The overall effect size was d = 0.33, 95% CI [0.09, 0.57].

0.75, 95% CI = [0.54, 0.95], t = 7.47, p < .001).

as ONE row: test_type = "t", stat_value = 7.47 from column B, effect_reported = 0.33 from column A, graded g_ind = 0.379 against N = 1555, status WARN. Both numbers appear in the article. The pairing does not. The 0.75 that belongs with t = 7.47 was dropped, and a reported effect was charged against a statistic it was never reported with. This was disclosed as a known limitation in the 0.6.20 entry; it is now fixed. The two values are separate rows — d = 0.33 with its interval as a d_reported_only row, t = 7.47 on its own as NOTE — and the orphaned effect is verified to survive, since losing it would be worse than the mis-pairing.

The fix is a lookahead, not a paragraph rule, and that distinction was decided by measurement rather than taste. The first version dropped the sentence anchor and split on any blank line followed by a capital or a digit — the shape a merge takes when the extractor truncates a column mid-sentence. Two independent cross-model reviews converged on the same class of counterexample, and the whole-corpus diff found the class in the wild:

input what the general rule did
..., p = .026, d = 0.75⏎⏎95% CI [0.09, 1.41]. t row loses its CI
..., p = .026⏎⏎Cohen's d = 0.75, ... t row loses its effect size — the consistency check silently stops running while the row still reports OK
The effect was medium, d = 0.⏎⏎65, 95% CI ... boundary inside a decimal
t(58) = 3.45, p = .03⏎⏎1, ... p published as .03 instead of .031
..., d = 0.65, 95%⏎⏎CI [0.40, 0.90], ... CI severed from its own effect
collabra.37122’s flattened appendix table OR = 0.99, 95% CI [0.77, 1.27] dropped from the output entirely

Requiring the boundary to sit after sentence-ending punctuation refuses every one of them: in each case the character before the blank line is a digit, a %, or a comma — text that is mid-statement by construction. Five reproduced as row-level defects and are pinned as regression tests watched RED against that first version; two did not reproduce (both sample-size cases, where the truncated N is rejected and the document-level scan supplies the value regardless) and are recorded as such rather than promoted.

(?<!\d\.) is the one guard the anchor does not supply: a period preceded by a digit is a decimal point, not a sentence end. Rejoining <digits>.⏎<digits> was considered and refused on corpus evidence — across the 48 real-article texts that shape occurs 221 times and is dominated by reference numbering and section headings (...962-967.⏎⏎30.⏎⏎Singh H, ..., which a joiner would fuse into 967.30), while all 22 occurrences of <label> = <digits>. at end of line are sentence-final periods. A bridge with no observed benefit and a demonstrated corruption path is the v0.6.20 defect class exactly.

Whole-corpus diff (48 real-article texts, 765 rows, per-row field comparison — not row counts, per the v0.6.20 lesson that nrow() cannot see value corruption): the general rule touched 11 files with 6 rows lost, 13 gained and 16 changed. The shipped rule produces one change in the entire corpus — the target defect, and nothing else. 0 rows gained, 0 lost.

Still not fixed, deliberately: a column truncated without sentence-ending punctuation still merges. Covering it costs a reported value on real text, which is worse on a science tool than one unchecked row.

Three separator defects, from the review v0.7.3 never ran

v0.7.3 shipped without /escicheck-qa, /escicheck-review or a cross-model review. Running them found the following. Twelve reviewer findings were produced across two models; every one was run against the working tree before anything was changed, three were refuted outright, and the three below reached a published number.

A confidence interval was being destroyed on real published text. An interval written with no space after the comma — 95%CI=[7.944,11.984], which is how nathumbeh_replication_2025 writes all 11 of its intervals — matched the European full-notation rule as 7.944,11 and was rebuilt into 7944.11, leaving the text as 95%CI=[7944.11.984]. Both bounds came back NA and the row still reported status OK. The same clause written with a space parses correctly, which is exactly what made it invisible: the paper’s other rows look fine. The guard is structural and needs no locale inference — in genuine full notation the part after the comma is a terminal fraction, so it cannot be followed by another decimal point and more digits. Both lookaheads in the fix are load-bearing: (?!\.\d) alone is defeated by backtracking, which gives up a digit and matches 7.944,1 instead.

A European p-value was silently dropped. Rule D1 required at least one digit before the comma, but a continental paper omits the leading zero exactly as APA does: p = ,025 is p = .025, and p < ,001 is p < .001. Neither matched, so t(48) = 2,31, p = ,025, d = 0,74 published the row with p_reported = NA while the same clause’s t and d converted normally — and the fuller form carrying p < ,001 returned zero rows. New rule D1b keys on the value position (a comma directly after =, < or > with no space), not on the locale, so an English document carrying a stray continental value is repaired rather than dropped. Requiring no space is what keeps CI [0.45, 0.89] and F(1, 30) out.

An OCR-spaced sample size lost everything after its first group. The repair loop’s anchor was \d{1,3}, so it could only ever run once: after the first join the prefix is four digits and the pattern stopped matching itself. nobs = 1, 234, 567 became nobs = 1234, 567 and the row published N = 1 234 — a wrong sample size, not a missing one.

Refuted, and recorded rather than “fixed”: that F (1,234) with an extracted space fuses its df pair (it normalizes to F (1, 234) and parses as df1 = 1, df2 = 234); that Indian grouping 12,34,567 is partially stripped (left verbatim); and that a 1:5 lopsided locale contradiction leaves d = 0,80 reading as 0 (it converts to 0.80).

Reproduced but deliberately left: F = 1,234 with no locale evidence reads as 1234 rather than 1.234 — genuinely undecidable from shape, and it publishes no effect size. RGB = 120,120,120 fuses to 120120120 — reproduced, and worse than the reviewer claimed, but the fused token is read as no statistic and the neighbouring row is byte-identical with and without it. A scan of the 48 real-article texts found 71 comma chains of three or more groups and every one is a citation superscript already protected upstream, so the residue is the known undecidable case rather than an observed defect.

The four normalization rules are part of the cross-language contract with docpluck, so all three fixes are added to inst/normalization-spec/ as conformance cases (spec 1.2.0 → 1.3.0, 61 → 69 cases) rather than only to the R implementation.

Measured, not implemented

Two items carried in the v0.7.3 handoff as open work were resolved by measuring the premise instead of writing code, and are recorded here so the measurement is not re-derived:

Frontend

Four test types added in v0.7.0 — wts, ats, brunner_munzel, yuen — shipped with no user-facing explanation, so the audit table named them and said nothing about what they are. The result-contract test had been failing since v0.7.0; npm run lint was never run for that release.

effectcheck 0.7.3

A comma between digits means at least four different things, and the code only knew two. v0.7.2 replaced the thousands/decimal whitelist with a shared spec. Testing that against the real 38-paper validation corpus — 2,984 comma occurrences, which the tmp/ corpus used until then contained none of — showed the remaining classes were larger than the ones already fixed.

Contextual signals, measured rather than assumed. A space after the comma is strong evidence against a thousands separator: 68% of no-space occurrences are the thousands shape versus 22% of spaced ones. The spaced+3-digit class is overwhelmingly reference-list furniture (Psychol. Methods 3, 424-453, Nature genetics 54, 437-449) and RGB triples. An earlier draft admitted a space to repair N = 1, 234, and on real strings that fused volume into page (3424-453), collapsed RGB = 120, 120, 120 into 120120120, and turned the table row All places 403,669 107,081 into 403.669107.081. The general rule now refuses a space; a narrow N =/nobs = rule repairs the count case, where the spacing genuinely is extraction damage.

Chains resolved by group width. Of 35 bare comma chains in the corpus, “every group after the first is exactly three digits” classifies 34 correctly: 1,000,000 and 1,054,908 are numbers; Figures 6,7,8, 10,14,19, Ye1,2,3,4 and 2017,10,39 (a year followed by citation superscripts) are lists. Previously the decimal rule converted only the first pair, producing 1.2,3,4,5 — worse than either answer.

Structural spans are lifted out before any numeric rule runs: head-noun lists, tuples of any length, bracketed indices, coded variables, DOIs. Self-identifying full notation (1.234,56, 1,234.56, apostrophe and NBSP variants) is resolved without locale knowledge. Terminator whitelists are replaced by a negative (?!\d) boundary — that whitelist had been patched for /, then %, then -, each after a silent failure.

Document-level locale inference (infer_numeric_locale) resolves the one shape structure cannot: 1,234. The two conventions are mutually exclusive, so a single unambiguous token settles a document. Markers are operator-guarded — a bare 0,1/0.1 misfires on coded variables, version numbers, ratios and lists (5 of 10 realistic strings); requiring a comparison operator fixes all five. Conflicting evidence is reported as conflict, never majority-voted.

Impossible values can no longer exit as OK. Cramér’s V, |r|, |rank-biserial|, |Cliff’s δ|, η², ω², R² are bounded by construction, and U/W are integers by construction. A violation is now WARN + extraction_suspect, with the value kept visible so a reader can recognise the failure. This catches the whole class independently of cause: it would have caught all three separator defects on its own.

Dual-p reporting. resampling_inference is a CLAUSE-level fact; whether the BOUND p is resampling-derived is a VALUE-level fact. Conflating them shipped a false claim — for the real sentence t(2037) = -3.26, P = 0.001, P-permutation = 0.002 the bound p is the parametric 0.001 (verified: 2*pt(-3.26, 2037) = 0.001132), yet the row asserted “this p-value is not reproducible even with the raw data” about it. New p_reported_is_resampling keys every claim about the bound value. The safety condition is provability: only a GLUED qualifier (P-permutation) is invisible to pat_p, so the bound value is provably parametric and is verified normally; a SPACED qualifier (permutation p = .062) may itself have been bound, so the row stays conservative. Corpus evidence: 13 real dual-p instances, all verified against pt(), of which 2 (15%) disagree about significance at α = .05 — routine, not a corner case. Magnitude thresholds remain refused, now empirically: agreeing pairs range from nearly identical to 29× apart, and discordant pairs are not magnitude outliers.

Also: a permutation result no longer worsens a verdict. Five rows in a PNAS paper reporting both p-values went OK → WARN purely for containing the word “permutation”, because the p-consistency block is also the WARN → OK rescue.

Suite 3319 passing / 0 failures across 135 files; conformance corpus 61/61. Validation corpus (25 papers, 550 rows): 0 rows lost, 0 value changes, 10 new extractions, 2 verdict changes. Two corrected values are the defect caught in published work — U = 55,890 and U = 84,716 in an RSOS paper were being read as 55.89 and 84.716.

effectcheck 0.7.2

A normalization rule was turning thousands separators into decimal points. normalize_text() resolved the decimal-comma / thousands-separator ambiguity with a whitelist of four syntactic contexts. Everything outside that list was corrupted:

input produced consequence
U = 12,345 12.345 rank-biserial 0.99938 published where the truth is 0.38275 — status OK
nobs = 1,182 1.182 → N = 1 Cramér’s V = 3.5355 published — V is bounded in [0,1], so that value is impossible — status OK
BF10 = 1,234,567.89 1.234,567.89 unparseable; a degraded BF flips “decisive” to “negligible”
1,234/5,678 arm counts NA risk-ratio verification silently skipped
M = 1,234.56, SE = 1,234.5, AIC = 12,345.6, H(2) = 1,234.56 two decimal points unparseable

The whitelist had already failed three times — downstream E8 (47 rows lost in one article), the N = case, and a resample-count case in v0.6.22 — each fix adding one more entry. It was an enumeration, not a rule, and nobs (a documented sample-size token since v0.5.5) was never in it.

Root cause. Two decimal rules allowed unlimited digits before the comma ((\d{1,3}), and ([-+]?\d+),). A European decimal has exactly one integer digit (0,05, 1,5, 9,81); two or more digits before a comma is a thousands group. That single constraint is the whole fix.

The rules now live in one place. inst/normalization-spec/ holds a versioned SPEC.md and a language-neutral conformance.json that both effectcheck and docpluck must satisfy. docpluck had already solved this correctly — its A3a rule handles every case above, and its source still carries the comment “ESCImate Request 1.1” recording that we asked for the fix and then never retired our own broken copy. Two implementations with no shared test cannot stay in sync; the corpus is that test. normalization_spec_version() is recorded so a published number carries the provenance of the rules that produced it.

We also found a case where docpluck is wrong, and filed it back (REQUEST_TO_DOCPLUCK_normalization_spec.md): its lookahead omits ,, so t(28) = 2,21, d = 0,45 leaves 2,21 unconverted and a parser reads 2 — the same class of silent error, in the other direction. A European decimal is routinely followed by the list comma. Both divergence cases are tagged in the corpus.

Deliberately kept as a separate layer: t(1,197) and F(7,140) are the same shape with opposite answers — a t-test takes one df (so 1197 is a thousands separator), an F-test takes two (so 7,140 is a genuine pair). That depends on test arity, which is statistics-aware, so the shared spec protects all X(…) brackets and effectcheck strips the single-df ones afterwards. docpluck cannot make that call and should not try.

Three defects surfaced during verification that the unit tests alone would have missed: the corpus caught my own wrong expectation about t(1,197); the end-to-end check caught / missing from the boundary set, which silently truncated a clinical arm count from 1234 to 234; and the corpus-diff harness was picking up this session’s scratch prompt files and reporting them as regressions.

Suite 3181 passing / 0 failures across 135 files. Whole-corpus render diff over 8 real articles / 169 rows: byte-identical. Note the corpus contains zero thousands-separated numbers, so it is structurally blind to this defect class — the conformance corpus exists precisely because the render diff cannot see it.

effectcheck 0.7.1

A rank test’s estimand, stated. The methodological point underlying the whole 0.6.21–0.7.1 line of work: switching from a t-test to Mann-Whitney when normality looks doubtful is not a like-for-like substitution. MWU targets stochastic superiority, P(X>Y) + 0.5·P(X=Y) — not a difference in means, and not generally a difference in medians either, since that reading requires a location-shift / equal-shape assumption papers rarely state.

The claim is exact, not rhetorical, and the test pins the arithmetic: X uniform on {1, 5, 6} and Y uniform on {4, 5, 9} have identical medians (5), yet P(X<Y) + 0.5·P(X=Y) = 11/18 = .611 — far from the .5 null, so the procedure has power against distributions whose medians coincide. (Fay & Proschan 2010, Statistics Surveys 4:1–39; Divine, Norton, Baron & Juarez-Colunga 2018, The American Statistician 72:278–286, “The Wilcoxon-Mann-Whitney Procedure Fails as a Test of Medians”.)

A NOTE, never an error, on U and W rows only. We deliberately do not scan surrounding prose for mean/median language: an interpretation sentence cannot be reliably linked to a specific test, and a false accusation there would be worse than silence.

Implementation note: written as a standalone block rather than a branch of the computation chain. Folding it in as } else if (tt %in% c("U","W")) terminated that chain early, which would have let its trailing clauses fire for unrelated test types.

Second cross-model review round (v0.6.22 / v0.7.0 / v0.7.1)

Eight further defects raised against the new material, all eight reproduced locally before being acted on. Four wrote a wrong number rather than merely a wrong flag:

Plus: B = 10,000 with no unit noun was missed; a strict inequality at the floor (p < .001 with 999 permutations) was not flagged because .001 < .001 is false, now compared with <=; a Monte Carlo SE and 95% interval were printed around an inequality p (p < .05 → “SE = 0.00218, interval 0.0457 to 0.0543”), attaching concrete numbers to a value the paper never reported; and — the commonest reporting form of all — Yuen and Brunner-Munzel written with a plain t(df) were claimed by the generic t branch and had ordinary Cohen’s d variants computed for them.

Suite 3024 passing / 0 failures across 134 files. Whole-corpus render diff: byte-identical at every stage of this work.

effectcheck 0.7.0

The modern nonparametric / robust family: Brunner-Munzel, ATS, WTS, Yuen. The reviewer asked whether the permutation work was “extendable to other tests (e.g. Brunner-Munzel, ATS, WTS etc)”, assuming that without raw data none could be reproduced. It is extendable, and further than the question supposed: all four report a statistic against a known reference distribution, so the reported p IS independently verifiable from the statistic and df alone. It is the effect size that is not recoverable — the same shape as the existing cochran_q branch (v0.5.15).

new test_type reference distribution p verification
wts chi-square(df), df = rank of contrast matrix pchisq(WTS, df, lower=FALSE)
ats F(df1, df2), df1 typically non-integer, df2 may be Inf pf(ATS, df1, df2, lower=FALSE)
brunner_munzel t(Satterthwaite df) 2*pt(-abs(W), df)
yuen t(trimmed df) 2*pt(-abs(t), df)

pf(F, df1, Inf) reduces exactly to pchisq(df1*F, df1); the identity is asserted in the test rather than assumed. Brunner-Munzel and Yuen require their name in the clause — their statistics are written W / W_BF / t, which collide with Wilcoxon and with an ordinary t-test, so the dispatch uses the same discipline chisq_subtype applies to McNemar/Friedman. All four branches sit before the generic ones, and a regression test pins that ordinary t / F / W rows are unaffected.

Each is exempted by the v0.6.21 resampling machinery when its own clause says permuted/bootstrapped — GFD reports both an asymptotic and a permuted WTS precisely because the asymptotic one is liberal in small samples, and bootstrap Yuen is standard in WRS2. Grading those against the parametric reference would be the v0.6.21 defect reintroduced through a new door.

Orientation warning, recorded because we got it wrong first. Brunner-Munzel’s estimand p̂ = P(X<Y) + 0.5·P(X=Y) is the complement of Vargha-Delaney’s A_XY (they sum to exactly 1), so δ = 1 − 2p̂, not 2p̂ − 1. Our design document had it backwards; a third verification pass caught it, and direct computation over three cases (including one with a tie) confirmed the sign flip. Shipping the original would have published every Brunner-Munzel effect with the wrong sign. No cross-check against a reported Cliff’s δ or Vargha-Delaney A is performed until the reporting orientation can be pinned.

Pre-existing defect fixed in the same run: .friendly_test_name() covered only 12 of the 25 shipped test types, so spearman, kendall, cochran_q, RR, rdpct, md_hl, binomial, interaction_p, mediation_indirect, mcnemar_or, bayes_factor, hazard_ratio and d_reported_only fell through to their raw slug in every report. All filled in, and a test now asserts every type in the default stats vector has a display name — so the next new type cannot silently regress it.

effectcheck 0.6.22

What CAN be checked on a resampling p-value without the raw data. 0.6.21 stopped grading a permutation p against the parametric reference. That left the reviewer’s actual question open: if it cannot be recomputed, can anything be verified? Three things can, none of which need the data.

1. The minimum attainable p. A Monte Carlo permutation p sits on the lattice (r+1)/(B+1) (Phipson & Smyth 2010), so with B stated the smallest value the procedure can produce by counting is 1/(B+1). "1,000 permutations, p < .0001" is below that floor. New resampling_B column parses the resample count.

2. The exact-permutation floor. With n1 and n2 known the reference set has choose(n1+n2, n1) members, so no exact p below 1/M is reachable — at n1 = n2 = 5, M = 252 and nothing under 1/252 exists.

3. Monte Carlo fragility. SE(p̂) = sqrt(p(1−p)/B), so p = .048 with B = 1,000 has SE ≈ .0068 and an approximate 95% interval straddling .05 — the significance decision is not stable at that resample count. Reported only when the interval actually crosses alpha.

Plus a reporting-completeness note: a resampling result that never states B is not reproducible even with the raw data.

None of these is a hard ERROR, deliberately. Legitimate practice reaches below the counting floor — GPD tail approximation, sequential Monte Carlo, combining per-stratum p-values, mid-p (0.5/M), randomized p (no positive floor at all). They state an arithmetic fact and let the reader judge.

Two calibration findings, both from enumeration rather than assertion:

New columns: resampling_B, resampling_p_below_floor. 8 test_that blocks, watched to FAIL first (11 failures + 2 errors before implementation). Suite 2947 passing / 0 failures across 132 files. Whole-corpus render diff: byte-identical, 0 rows gained, 0 lost, 0 verdict changes.

effectcheck 0.6.21

A resampling-derived p-value was being graded against a parametric reference. Raised by a methodologist reviewing ESCImate, who asked whether a “perm Welch t” could be approximately checked. Investigating it surfaced two defects, both reproduced at 0.6.20 before any fix was written.

The false flag. "A permutation Welch t-test with 10,000 permutations showed no significant difference, t(58) = 2.31, p = .062" returned decision_error = TRUE, reason = reported_ns_computed_sig. The paper is correct: the parametric p for t(58) = 2.31 is .0245, but the reported .062 came from the permutation distribution. The pre-existing method_context_in_chunk cap did not rescue it — “permutation” was absent from method_kw, and that cap only fires on status == "ERROR", not WARN.

The wrong published interval. ci_OR_all() back-derives a CI from the reported p via SE = |log(OR)| / qnorm(1 - p/2), an inversion that assumes a normal reference distribution. Fed a permutation p, a McNemar row reporting OR = 2.50, 95% CI [1.05, 5.95], p = .062 produced a computed interval of [0.9551, 6.5441] — crossing 1 where the paper’s does not — and then declared the paper’s own correct CI INCONSISTENT. Reachable from four test types (mcnemar_or, chisq, regression, z); it requires a reported CI for the comparison to fire, which is why a first probe without one returned NA and looked clean.

The fix is scoped to the p-value, not the row. A permutation changes only the reference distribution — the statistic itself is computed identically, so d = 2t/sqrt(df) remains exactly as valid and is still checked. Blanket-capping the row at NOTE would have discarded a real, correct check. New parse columns resampling_inference / resampling_method read the row’s own clause (the v0.6.18 Welch precedent — reading context_window leaked a modifier onto a neighbouring row, N 132 → 403). When set: the p-consistency comparison and decision_error are suppressed, the p-back-derivation is gated off, and an honest note names the parametric comparator without asserting an error.

Deliberately not a numeric |p_perm − p_param| threshold. Under the very conditions that motivate permuting, the two legitimately differ, so any magnitude threshold would manufacture false flags; only decision discordance is reported, and only as a note. Verified by simulation (1,000 sims, equal means, nominal α = .05): permuting the raw mean difference gives 26.6% type-I error at n₁=10 (sd 3) vs n₂=40 (sd 1), while permuting the Welch t itself gives 7.2% against the parametric 6.4% — so a reported “permutation Welch t” must mean the studentized form (Janssen 1997), and its p sits close to, but not on, the parametric p.

Two deliberate exclusions in the keyword set: "randomization" must be qualified by "test" (a bare randomi[sz]ed would match every randomized controlled trial), and "exact test" is absent entirely — Fisher’s exact is a closed-form conditional test whose p is computable, so matching it would suppress a legitimate check. Both are pinned by tests, as are "randomly assigned" and "randomized controlled trial".

Also: "bootstrap" belongs to both method_kw and the resampling set, so a bootstrapped result used to be explained as a methods-section artifact (“power analysis, meta-analysis, etc.”) — a false statement about the row. The method-context message is now suppressed when the row is recognised as resampling-based.

Cross-model review raised eight further paths against the first draft, and all eight reproduced locally before being acted on — three over-suppression, four under-application, one regex gap:

14 test_that blocks, every defect-targeting one watched to FAIL against the unfixed code first (23 failures + 1 error at clean HEAD). Two blocks are green at HEAD by design — they guard against this fix over-reaching. Suite 2911 passing / 0 failures across 131 files (baseline at HEAD 2849). Whole-corpus render diff over 9 real-article texts, 192 rows: byte-identical, 0 rows gained, 0 lost, 0 verdict changes.

effectcheck 0.6.20

A normalization rule was deleting reported statistics. downstream filed two apparently unrelated defects (O-1, O-2) with two different diagnoses. Both traced to one line in normalize_text(), and neither diagnosis was right. The sweep that followed found four more defects of the same class — one of them destroying a real effect size in a paper already in the regression corpus.

The defect class: normalization that deletes source text

normalize_text() repairs PDF text-layer artifacts before anything is parsed. Several of its rules “bridge” a line wrap by skipping a span and adopting a number from the next line — the repair for d =\n0.80. None of the three that skipped a bounded span checked whether that span contained a value, so a reported statistic was destroyed and replaced by whatever number happened to open the next line:

"etap2 = .86, and Experiment\n1b"                 -> "partial eta-squared = 1"
"d = 0.74 (see Table\n2)"                         -> "p = 2)"
"chi2 (4, n = 211) = 12.74, p = .013\n\n10 items" -> "chi2 (4, n = 10 items"
"t(48) = 2.31, p = .025, d = 0.65\n\n10 items"    -> "t(48) = 2.31, p = 10"
"r(351) = .164, p = .050\n\n10 items"             -> "r(351) = .164, p = 10"

Two invariants now govern every bridging rule:

  1. The skipped span must be digit-free. A bridge may never discard a number that is already present. (One sibling rule already did this and was correct; the difference had gone unnoticed for eleven releases.)
  2. When prose is skipped, the adopted number must carry a decimal point. A bare integer opening a line is a page or section number, not a statistic. The genuinely bare case — an assignment whose line simply ends at the =, where a wrapped integer n =\n120 is legitimate — is handled by the whitespace-only joiner, which deletes nothing.

Corrections to the filed diagnoses, since downstream is writing this up:

Four more defects of the same class, found by the sweep

downstream O-3, O-4, O-5

Also

Two defects in this release’s own diff, caught by the release pipeline

Recorded because both would have shipped an inaccuracy, and both were found after the change was otherwise green (suite passing, CRAN check clean, corpus diff clean, two cross-model reviews complete).

Verification

Known limitation (new, deliberate)

The sentence splitter does not break on a paragraph boundary, so a two-column merge like the spps.txt case above still yields one row pairing an effect from one column with a test statistic from another. Splitting on a blank line would separate them (and the bare-d-with-CI pattern would pick up the orphaned effect), but it changes chunking for every document and belongs in its own release with its own corpus validation.

effectcheck 0.6.19

Two shipped “cannot verify” messages claimed mathematical impossibility that does not hold. Found by the 2026-08-05 re-audit of every won’t-fix / not-recoverable ruling in the repo, run under the portfolio-wide triple-verification rule (each claim checked against the primary source, then codex, then a Sonnet pass instructed to refute). No verdict, status, or computed value changes — the conservative behaviour was and remains correct. Only the stated reason was wrong, and on a science platform a false claim about what is mathematically knowable is itself a defect.

Regression tests: the two message assertions previously grepped the retracted phrases, so correcting them would have looked like a regression. Both now pin behaviour plus semantic content and add an explicit expect_false on the retracted phrase — verified RED against the old wording before being restored.

Full report: communications/FINDINGS_2026-08-05_rejection_reaudit.md.

effectcheck 0.6.18

Four sample-size defects that published wrong or unlabeled Ns, found by the 2026-08-04 escicheck-iterate canary audits of collabra.57785, collabra.90203 and pci.rr.100726. Every one was reproduced at HEAD before any code changed, and each carries a regression test that was watched RED against the unfixed code.

A post-hoc contrast that reprints the omnibus ANOVA error df is no longer scored against df + 2. After a k-level ANOVA, stats packages routinely reprint the omnibus error df on each pairwise contrast, but a two-level contrast uses only ~2/k of the sample. On collabra.90203 an F(2, 998) omnibus (N = 1001) is followed by t(998) = 2.46, p = .041, d = 0.19 [0.04, 0.35]; binding N = 1000 computed d = 0.1556 (delta 0.0344) and produced a WARN plus an INCONSISTENT CI flag, both false — at the true contrast N = 669 the computed d = 0.1902 (delta 0.0002) and the CI reproduces the reported one.

The rule ships in a deliberately conservative form. A cross-model review (Codex) refuted the first draft, which fired on same-document df equality alone: a paper can legitimately contain a 3-arm F(2, 998) and a real two-group comparison of the full sample with a genuine t(998). That counterexample was reproduced locally before the design changed. So omnibus-df matching now only proposes a candidate N, and adoption requires BOTH that the surrounding text describes a post-hoc / pairwise / Tukey / Bonferroni comparison AND that the row’s own reported effect size is explained materially better by the candidate than by df + 2.

A second review round (also Codex) showed why the text requirement is load-bearing: effect-size fit alone is circular, because an unequal-groups full-sample t reports a d larger than the equal-n df + 2 estimate and can fit the candidate coincidentally — a document with F(2, 98) and a genuine t(98) = 2.00 whose unprinted cells are n1 = 20, n2 = 80 would have been rewritten to N = 67 against a true 100. CI-width corroboration does not rescue it (unequal groups widen the interval in the same direction a smaller N does), and contrast_N / (df + 2) is 2/k in both cases, so the two are mathematically indistinguishable from the numbers alone. Both counterexamples are pinned as tests.

A row whose reported effect matches neither N keeps its inconsistency flag; suppressing it unconditionally would mask a genuine reporting error. Explicit per-group sizes always outrank the hypothesis, and a 2-group F(1, df) never triggers it. Adopted Ns are labeled N_source = "omnibus_df_contrast" and disclose the balanced-cell assumption in uncertainty_reasons.

Verified on the real corpus render: the two false positives clear (WARN → PASS, CI INCONSISTENT → MATCH) while the same paper’s honestly-reported t(667) / t(668) rows stay byte-identical.

A cross-model review of this release’s own diff surfaced six further defects, each reproduced locally before being fixed: subgroup_sum was exempted from the df-authority override although it is matched over the wider context (so a subgroup pair in [df+3, df+12] bypassed it); the Welch branch’s implausible-N cross-check was still keyed on global_text alone, which became the only remaining guard once stated-Welch rows began skipping the non-Welch path; the Welch back-computation N = 4t²/d² treated a reported Hedges’ g as a Cohen’s d (now converted via d = g / J, after a first attempt that simply excluded g proved worse — it left the implausible scraped N in place); the new relative p-value gate falsely flagged legitimate rounding (p = .01 printed at two decimals honestly represents a computed .0051), so the ratio now applies only outside the rounding band the printed precision implies; and N_source = "df_inferred" was not emitted when a scraped source had been discarded and replaced, leaving the row advertising a provenance its published number no longer had.

The z-branch scraped-N disclosure now covers every scraped source. v0.6.17 added the warning because a z-test has no df, so none of the df-keyed N-plausibility guards that protect t-rows can fire — whatever N is bound is used for d = 2z/sqrt(N), dz = z/sqrt(N) and r = z/sqrt(z² + N) with nothing to contradict it. But it was keyed on global_text alone, leaving local_context / extended_context / subgroup_sum silent. "The calibration sample (N = 100) was used first. In the target subsample (n = 25), z = 2.00, p = .046, r = .20." bound N = 100, published status = "OK" with an entirely empty uncertainty_reasons, and computed r = 0.196 against the reported .20 — an apparent match, where the clause’s own n = 25 gives 0.371. Re-keyed on .SCRAPED_N_SOURCES. Found by the final pre-push cross-model review.

Known, unfixed: a statistic quoted twice in running text is still scored twice. pci.rr.100726 is a peer-review letter whose comment prints one t(868) = -3.01, p = .006 twice to illustrate APA comma placement, and the render emits two rows. Three dedup rules were built and each was disproved — by the v0.6.14 invariant that two genuinely distinct correlations can share every reported number (r(797) = .16 twice in one sentence, different variables), and finally by the real paper itself, whose echo carries its own t(df) anchor and so is indistinguishable from a second result. Separating the two needs a signal this layer does not have. The duplicate is left in place deliberately: a duplicate row is a counting error the reader can see, a dropped row is a lost result they cannot. The invariants any future fix must respect are pinned in test-v0618-prose-restatement-dedup.R.

A reported CI with no parseable effect size is no longer silent. Such a row narrowed to a p-value-only check and could still publish status = "OK" with ci_check_status = "MATCH" and nothing in uncertainty_reasons — a reader saw a green row and could not tell the paper had reported an effect size the tool never verified. (On collabra.90203 the partial-eta-squared symbol is drawn in the source PDF as filled vector curves with no character object at all, so the body text arrives as a nameless = .008; the value is recovered from the table view, but silence on the body-text row was still dishonest. Mechanism corrected 2026-08-05: this entry originally said the glyph “has no ToUnicode mapping” — that cause is refuted; the symbol is ink, not badly-encoded text. The behaviour described here is unchanged and correct.) A confidence interval cannot exist without an estimate, so its presence is proof an effect size was reported — the row now says the effect size was not verified. Found by the cycle-2 canary audit.

Also fixed in the same render: the ANOVA-design uncertainty message was authored with an escaped em-dash, which passed the source-file ASCII check but reached the user as a corrupted byte. A new test asserts emitted messages contain no non-ASCII bytes — checking the source alone was blind to this.

Internal: pat_N is hoisted to a package-level .pat_doc_N with a shared .doc_global_n() helper, so check_text() and parse_text() cannot drift apart (an attribute on parse_text()’s return value was tried first and silently vanished on the zero-statistics early-return paths).

effectcheck 0.6.17

Three sample-size defects that published wrong effect sizes, found by expanding the escicheck-iterate comparison harness onto papers that had an AI stats gold but had never been compared against the library (~548 gold results the library had never been audited against).

Also verified on the fixed-3 canary: 10.1525/collabra.90203’s t(998) pairwise contrasts now bind N = 1000 (df+2, the two conditions actually compared) instead of the paper-level 1004, and their reported-vs-computed CI mismatch shrinks accordingly.

effectcheck 0.6.16

Nine audit findings from the 2026-08-04 canary sweep (CI sign-alignment E3, honest design labels E2, self-consistent deltas E4, Cochran-Q sample-size guard E5, multiplicity-adjusted-p guard E6, own-clause N binding E7, reproducible repro-code E8, and two recovered PARSE-MISS classes E10/E11). Two further audit findings were REFUTED by local reproduction and documented rather than “fixed” (see communications/TRIAGE_iterate_2026-08-03.md).

Earlier findings from the same sweep: CI sign-alignment (E3), an honest design label for bare table rows (E2), a self-consistent reported-vs-computed delta (E4), a Cochran-Q sample-size guard (E5), and a multiplicity-adjusted-p decision-error guard (E6).

CI comparison follows the magnitude convention of the value match (sign-alignment).

effectcheck 0.6.15

Restatement guard for the v0.6.14 prose-dedup un-collapse + Mode B typed-n binding.

effectcheck 0.6.14

Correlation deduplication correctness fix.

effectcheck 0.6.13

Two new test types + two canary-re-audit fixes from the 2026-07-02 escicheck-iterate cycles 2-3 (F1 bare Bayes factor + HR hazard ratio, both user-approved; independent Sonnet-watches-Opus canary re-audit over the fixed-3 + rotating set).

Two new N_source values (own_clause_arms) and two new test_types (bayes_factor, hazard_ratio) are documented in API.md. Cycle-3’s HR feature is orthogonal to the canary papers: 4 of the 5 canary renders were byte-IDENTICAL to their cycle-2 PASS renders (a deterministic diff carries the prior verdict), and cog_emo was re-audited PASS. Full suite 964 test_that blocks / 0 fail, R CMD check --as-cran 0E/0W. Residual canary findings are all docpluck-boundary and filed to the 2026-07-02 extraction-tool defect log (DP-4 collabra.37122 loc-202 figure-caption CI truncation; DP-5 s41598 shredded survival table; DP-6 cog_emo garbled Table-7 duplicate; DP-3 collabra.57785 Table-8 Importance d/CI+design re-confirmed).

effectcheck 0.6.12

Three fixes from the 2026-07-02 escicheck-iterate cycle-1 canary re-audit (independent Sonnet-watches-Opus over the v0.6.11 canary set), all on collabra.57785 (Experiential-vs-Material Purchases replication+extension of Carter & Gilovich 2012).

Full suite 924 test_that blocks / 0 fail; R CMD check --as-cran 0E/0W. Regression tests in tests/testthat/test-v0612-ownclause-n-and-repcol-dedup.R. Two docpluck text-extraction defects were filed (NOT effectcheck defects) to the 2026-07-02 extraction-tool defect log: DP-1 (collabra.77859 camelot_t10 Study-1 Table-1 binds the wrong column as t/d/df/CI — delivered t = 0.6 where the gold reads t = 5.65, so effectcheck faithfully rendered docpluck’s wrong values), DP-2 (collabra.77859 “Expensive” manipulation-check row t = 15.57 not delivered as a flattened_row). A standalone-Bayes-factor gap on collabra.90203 (bare BF01 = 0.11 / 1.24 not extracted) was surfaced for a product decision rather than fixed — the paper reports 13 BF01 = values of which the gold wants only 2 as standalone results, and no parse-pattern rule reliably separates the 2 primary-analysis Bayes factors from the 11 supporting/companion ones (see communications/TRIAGE_iterate_2026-07-02.md F1).

effectcheck 0.6.11

Two fixes from the 2026-07-01 escicheck-iterate cycle-2 canary audit (independent Sonnet-watches-Opus over the v0.6.10 canary set).

Full suite 917 test_that blocks / 0 fail, R CMD check --as-cran 0E/0W. Regression tests in tests/testthat/test-v0611-origcol-and-mdhl-n.R. Surfaced alongside three new corpus golds generated via article-finder for the deeper audit (collabra.74820 neuroticism×EC moderation, collabra.122515 creativity-depression mediation, collabra.88158 daily-diary multilevel — the last dropped from the audit set due to a source-PDF binding defect: pages 7-12 are a mis-bound different article).

effectcheck 0.6.10

A bootstrapped mediation indirect effect reported with a Sobel Z is now a first-class test_type = "mediation_indirect", from the 2026-06-29 escicheck-iterate new-corpus pass against the Outcome Bias replication+extension (collabra.126266, Aiyer/Chan/Feldman 2024).

A clause like “the bootstrapped indirect effect of X on Y was .05, 95% CI [-.04, .12], Sobel Z = 0.84, p = .40, ACME found to be robust until ρ = 0.7” previously routed the Sobel Z = 0.84 to a PLAIN z-test, and then the fallback effect-size pattern grabbed the sensitivity-analysis ρ = 0.7 — the value of the error-term correlation at which the ACME mediation stops being robust (an Imai/Keele/Tingley sensitivity bound) — as the EFFECT SIZE, discarding the actual indirect effect (.05) and emitting a spurious WARN. All four mediation rows (H2 + H5) were mis-typed z with effect_reported_name = "rho" and effect_reported = the sensitivity bound.

parse.R adds pat_mediation_indirect (anchored on “indirect effect … was ” AND “Sobel Z = ”, both required) and a dispatch branch that classifies the row mediation_indirect, binds the indirect-effect coefficient as the reported effect (effect_reported_name = "indirect_effect"), the Sobel Z as the test statistic, and the bootstrapped CI as the indirect-effect CI (anchored at the indirect-effect value, before the trailing ρ). An is_mediation_indirect flag suppresses the fallback-ES ρ grab. check.R routes mediation_indirect to an honest extraction-only NOTE (the indirect effect is not recomputable from the reported numbers without the a/b path coefficients) and excludes it from the Phase-9 SKIP downgrade so the indirect effect + CI are surfaced. mediation_indirect is added to the check_text() stats allowlist and documented in API.md. The same new-corpus audit confirmed all 28 other reported statistics (replication/extension ANOVAs + Welch t post-hocs) are extracted correctly and the gold’s Table-5 “Original” (Gino 2009 comparison) rows and abstract-only effect-size restatements are correctly NOT extracted.

It also fixes the malformed p = <.001 form (a spurious = immediately before the real </> operator — a common PDF text-layer artifact) in pat_p / pat_p_sci / pat_p_enote: the operator group now accepts an optional leading = ONLY when a real </> follows (a lookahead), so p = <.001 parses to p < .001 while a normal p = .40 still captures = and p <= .05 still captures <=. Surfaced by the same collabra.126266 H5 punishment mediation row, where docpluck delivers “Sobel Z = 4.87, p = <.001” (the PDF prints “p < .001”) and the p had been dropped (p_valid = FALSE). Full suite 909 test_that blocks / 0 fail, R CMD check --as-cran 0E/0W. Regression tests in tests/testthat/test-v0610-mediation-indirect-sobel-z.R.

effectcheck 0.6.9

A =-as-U+00BC glyph-corruption normalization, from the 2026-06-29 escicheck-iterate new-corpus pass (SPPS “Inaction Inertia” replications, 10.1177/1948550619900570).

Some PDFs encode the = glyph such that the text layer emits U+00BC (“¼”, the fraction one-quarter). A whole paper can come through with EVERY equals sign as U+00BC and no real = at all (this SPPS paper: 120 U+00BC, zero =), so t ¼ -7.81, F (3, 1791) ¼ 200.12, d ¼ 0.57, M ¼ 20.20 all parsed to nothing — the entire body-prose statistics surface was invisible. normalize_text() now folds U+00BC → = ONLY in a statistical-operator position (flanked by whitespace and adjacent to a value / sign / bracket / a stat-word like “confidence”), so a genuine one-quarter fraction in prose (“¼ cup of sugar”, “¼ of participants”) is NOT rewritten. This is the same class of character-level normalization as the existing U+2212-minus and U+FFFD-eta-squared recovery. +11 results recovered on the SPPS paper; ZERO change on the canary + sweep corpus (they contain no U+00BC). Regression tests in tests/testthat/test-v068-equals-glyph-u00bc.R. The corruption is also filed to docpluck (the 2026-06-29 extraction-tool defect log §5) as the preferred upstream fix so all consumers benefit. Full suite 904 test_that blocks / 0 fail; R CMD check --as-cran 0E/0W.

The same SPPS new-corpus audit filed four docpluck table-extraction defects (sign-stripped negative table-cell t-values, a camelot_t11 t→F + d→p mis-typing of pairwise tests, an undelivered df column on Table-4 ANOVAs, and figure-embedded forest-plot estimates) — none of which are effectcheck defects; see the handoff §5.

effectcheck 0.6.8

Six parser/classification fixes from the 2026-06-29 escicheck-iterate canary audit (independent Sonnet-watches-Opus over the Collabra / PCI-RR / PLOS-Med canary set).

Full suite 901 test_that blocks / 0 fail; R CMD check --as-cran 0E/0W. Regression tests in tests/testthat/test-v068-*.R (6 files). One residual item filed (non-canary): collabra.23443 Table-5’s 4 one-sample-vs-mu=0 rows arrive as docpluck flattened rows whose one-sample design lives only in surrounding body prose (not on the row) — routed to the 2026-06-29 extraction-tool defect log (docpluck enhancement: carry the table’s introducing design onto the flattened row), alongside the docpluck table-shred / untyped-est handoffs.

effectcheck 0.6.7

Consumes the two newly-typed docpluck v2.4.98 table fields (fields.eta2 / fields.r), the reply to ESCImate’s 2026-06-25 docpluck handoff (DP-3 / DP-5). docpluck v2.4.98 now types the partial-η² column on a structurally-identified ANOVA table (fields.eta2) and types correlation-matrix cells (fields.r with rejoined CIs). Verified by re-extracting collabra.90203 + cog_emo from live docpluck v2.4.98 (/api/version → library.version 2.4.98) and checking against the AI stats gold. Full suite 877 test_that blocks / 0 fail, R CMD check --as-cran 0E/0W. Regression tests in tests/testthat/test-v067-docpluck-v2498-eta2-r.R.

effectcheck 0.6.6

Six parser/classification fixes from the 2026-06-25 escicheck-iterate canary audit (double independent Sonnet audit over the 5-paper canary set, after reconciling the 2026-06-23 Dropbox ai_gold conflict and regenerating the collabra.90203 + cog_emo stats golds). Verified against the AI stats gold; full suite 868 test_that blocks / 0 fail, R CMD check --as-cran 0E/0W.

effectcheck 0.6.5

Five parser/classification fixes from the 2026-06-21 escicheck-iterate canary audit (7-paper Collabra/RR/PLOS-Med set, independent Sonnet verification of the docpluck v2.4.95 production path). All verified against the AI stats gold; full suite 2117 pass / 0 fail.

Regression tests: test-v065-mcnemar-subtype-guard.R, test-v065-welch-not-paired.R, test-v065-beta-precedes-t-binding.R, test-v065-bare-binomial.R, test-v065-chi-subtype-gof-vs-independence.R. The canary-audit harness scripts/render_for_audit.R was also fixed to pass table_rows (mirroring the production /process path) so the audit exercises the v0.6.4 Mode B consumer.

effectcheck 0.6.4

Mode B docpluck table-row consumer (REQUEST_11 / docpluck v2.4.95). check_text() gains an optional table_rows argument that ingests docpluck’s structured flattened_rows[] (from POST /api/extract?structured=true, docpluck v2.4.95+) — the typed table-cell statistics that have no inline APA form in the prose. This captures the table-only results the 2026-06-16 canary audit flagged as PARSE-MISS (deferred in v0.6.3 because the hosted API previously column-shredded tables).

Regression tests in tests/testthat/test-v064-docpluck-table-rows.R.

effectcheck 0.6.3

Three fixes from the 2026-06-16 escicheck-iterate canary audit (communications/TRIAGE_iterate_2026-06-16.md):

effectcheck 0.6.2

Exact binomial test reported with Cohen’s h. New test_type = "binomial" matched via pat_binom_h, anchored on a “binomial p [op] ” clause followed (within ~80 non-period chars) by “Cohen(’s)? h = ”. When a “ out of ” clause is present in the same verbatim, N is recovered (N_source = "binom_n_out_of_N") and check.R re-computes the two-sided binomial p via stats::binom.test() assuming p_null = 0.5 (the most common null in binomial-vs-chance reporting); the recomputed vs reported delta appears in uncertainty_reasons. When N isn’t recoverable, status routes to NOTE – the Cohen’s h is accepted as reported.

Surfaced by the 2026-05-25 escicheck-iterate corpus expansion against the CRSP decoy-effect papers (Xiao/Zeng/Feldman 2021 et al), where 2-5 binomial-with-h rows previously fell through to WEAK_GOLD or OUT_OF_SCOPE. The NOTE-only template (LESSONS.md “NOTE-only test_type template”) was extended cleanly: parse layer adds the pattern + dispatch branch, check.R adds a tt == "binomial" branch with conditional recompute. A v0.6.3 follow-up could detect a stated null proportion (“vs 1/3 chance” etc.) to replace the p_null = 0.5 default.

Regression tests in tests/testthat/test-v062-binomial-h.R (7 cases: full CRSP verbatim with N recovery, bare binomial+h with N=NA NOTE, 80-char-lookahead far-apart rejection, “h” without “binomial p” anchor guard, chisq+h still routes to chisq, lowercase “cohen h” form, and uncertainty-message contents when N is recovered).

effectcheck 0.6.1

Bare t = X, p [op] Y (no df) extraction. Surfaced by the Lee-Feldman 2025 RSOS Newman-2014 RR replication during the 2026-05-25 escicheck-iterate corpus expansion (24 occurrences in one paper’s Tables 10-15: compact <label> M = m (sd), t = X, p < .001 form where df lives only in the table header, not the immediate sentence). Before v0.6.1 such reports returned 0 rows from check_text().

A new pat_t_p_nodf pattern matches t = X followed within ~80 chars by a p [<=>] clause; (?<![a-zA-Z]) keeps dt =, pt =, etc. from false-positive matching, and the 80-char lookahead bound prevents a stray t = X from being yoked to an unrelated downstream p = in long prose. df1 stays NA — check.R routes to status NOTE because the exact p-check needs df. Dispatch position: AFTER pat_t_nodf (t = X, df = Y form keeps priority and yields status=OK with full verification when df is present).

Regression tests in tests/testthat/test-v061-bare-t-p-nodf.R.

effectcheck 0.6.0

Clinical-trial RR / rdpct / md_hl independent verification, completing the v0.5.16/17/18 PROSECCO-trial test-type set. Closes the deferred v0.6.x follow-through promised in the v0.5.16-18 NEWS entries.

Verification (the v0.5.x NOTE rows now compute a comparison)

Regression tests in tests/testthat/test-v060-rr-rdpct-mdhl-verification.R. Closes the 2026-05-25-v06x-clinical-trial-compute-branches handoff.

effectcheck 0.5.18

Median-difference (Hodges-Lehmann) with IQR + CI (escicheck-iterate cycle 8). Completes the PLOS Med PROSECCO-trial PARSE-MISS punch-list opened in cycle 1.

New test type

effectcheck 0.5.17

Risk-difference percent with CI (escicheck-iterate cycle 7).

New test type

effectcheck 0.5.16

Clinical-trial risk ratio with two-proportion slash counts (escicheck-iterate cycle 7).

New test type

effectcheck 0.5.15

Cochran Q meta-analytic heterogeneity test (escicheck-iterate cycle-5, after user scope decision 2026-05-24 to bring Q in-scope).

New test type

effectcheck 0.5.14

Two narrow parse fixes from the 2026-05-24 escicheck-iterate cycle-4 validation against the Collabra canary.

Parse fixes

effectcheck 0.5.12

Recall fix for the Collabra / APA partial-eta-squared convention.

Parse fixes

effectcheck 0.5.11

Documentation-only release. The design_ambiguous output flag has always combined two semantically distinct cases under one name; this release makes the distinction explicit and parseable without changing behaviour.

Documentation / output-string clarifications

effectcheck 0.5.10

Bare r = with a confidence interval — a parse fix found by escicheck-iterate.

Bug fixes

effectcheck 0.5.9

Chi-square chi^2 caret token — a parse fix found by escicheck-iterate.

Bug fixes

effectcheck 0.5.8

Chi-square bare-n sample size — a parse fix found by escicheck-iterate.

Bug fixes

effectcheck 0.5.7

DSCF (Dwass-Steel-Critchlow-Fligner) post-hoc W — a parse + categorisation fix found by escicheck-iterate.

Bug fixes

effectcheck 0.5.6

Bare regression-coefficient lines — a parse fix found by escicheck-iterate.

Bug fixes

effectcheck 0.5.5

JASP “nobs” sample-size token — a parse fix found by escicheck-iterate running effectcheck against the real-article AI gold corpus.

Bug fixes

effectcheck 0.5.4

Regression-coefficient handling — a categorisation fix found by escicheck-iterate running effectcheck against the real-article AI gold corpus.

Bug fixes

effectcheck 0.5.3

Scientific-notation p-values — a parse fix found by escicheck-iterate running effectcheck against the real-article AI gold corpus.

Bug fixes

effectcheck 0.5.2

Subscripted chi-square notation — a parse fix found by escicheck-iterate running effectcheck against the real-article AI gold corpus.

Bug fixes

effectcheck 0.5.1

Stage 1 validation fixes — four gaps found by validating the v0.5.0 Stage 1 coverage against six real articles (AI gold generated via the article-finder skill).

Bug fixes

effectcheck 0.5.0

Coverage Stage 1 — closes effect-size / test-type gaps from the 2026-05-16 coverage roadmap (P1, P2, P3, P6, P7).

New features

Internal

effectcheck 0.4.2

Bug fixes

effectcheck 0.4.1

Bug fixes

effectcheck 0.4.0

Breaking changes — extraction layer removed

All file-input functions are now .Defunct() and emit an error directing callers to extract via docpluck and pass the resulting text to check_text():

The pure-text-analysis API (check_text(), compute_and_compare_one(), the parsing layer, all effect-size and CI computations, and every output column) is unchanged.

The package no longer requires poppler-utils, tesseract, magick, or qpdf system dependencies. SystemRequirements field removed from DESCRIPTION; corresponding entries removed from Suggests.

Migration: see https://docpluck.app/api-docs for the API contract. Working R reference implementation in the ESCImate web-app repo at tests/scripts/docpluck_shootout.R.

New features (carried over from 0.3.6 deception-detection work)

API documentation

effectcheck 0.3.5

Addresses downstream v0.3.5 request: CI-audit feature pack. Adds CI computation coverage for previously-uncomputable effect-size families (OR, R², standardized β, partial r, semi-partial r) and new per-row metadata for characterizing CI reporting quality at scale (precision tracking, completeness flags, level mismatch, bounded-parameter clipping, symmetry classification).

Purely additive — no v0.3.4 behavior changes.

Compute: CI computation coverage gaps closed

Parse: decimal-place precision tracking

Check: CI audit metadata (Phase 6)

Frontend (escimate.app)

effectcheck 0.3.4

Addresses downstream v0.3.4 request: 42 Category A ERROR false positives where reported eta2/etap2 was cross-matched to cohens_f/cohens_f2 without detection.

Check: Phase 8D Signal 14 — eta/f cross-family detection (E11)

effectcheck 0.3.3

Follow-up to 0.3.2 addressing downstream v0.3.3 request: the E8 pre-strip was a no-op on real docpluck output.

Parse: thousand-sep comma strip now handles spaces after comma (E8 follow-up)

effectcheck 0.3.2

Follow-up to 0.3.1 addressing downstream requests E8 and E10.

Parse: thousand-separator commas in test-statistic parens (E8, HIGH)

Compute: Cohen’s dz CI uses noncentral-t inversion (E10, MEDIUM)

Parse: decimal-comma no longer corrupts author affiliation markers

E9 — Smaller parse.R gaps (deferred, needs repro bundle)

effectcheck 0.3.1

This is a housekeeping release packaging the v0.3.0f → v0.3.0n bug-fix wave with a stable CRAN-style version number, batch-stdout hygiene, a schema stability test, and a new decision_error_reason diagnostic column. Addresses downstream requests E1–E4 and E7.

DESCRIPTION version sync (E2)

Batch stdout: noncentral-t overflow spam silenced (E1)

Schema stability test (E3)

New column: decision_error_reason (E7)

Expected row-count delta vs v0.3.0f (E4 — downstream batch guidance)

On the downstream downstream_regression 200-PDF frozen benchmark (seed 42), comparing v0.3.0f (last full batch) to v0.3.0n / 0.3.1:

subset v0.3.0f rows v0.3.0n rows delta v0.3.0f ERRORs v0.3.0n ERRORs
meta_psychology (139) 464 464 0 0 0
downstream_regression 2,209 3,385 +1,176 (+53%) 13 0

The +53% row-count delta on downstream_regression is driven by parser gains, not a config-default change (plausibility_filter and try_tables defaults are unchanged). The new rows come from:

Downstream consumers must re-derive all aggregate numbers from a fresh v0.3.1 batch — old v0.3.0f aggregates are not directly comparable. The 13 → 0 ERROR reduction on downstream_regression is real (v0.3.0n’s F ≈ 0 crash fix + multi-predictor-beta fix), not artefactual.

No columns were added or removed vs v0.3.0n other than the new decision_error_reason column described above.

effectcheck 0.3.0n

Bug fixes (downstream v0.3.0m batch deep-dive)

effectcheck 0.3.0m

Bug fixes (downstream batch validation)

effectcheck 0.3.0l

Enhancements

effectcheck 0.3.0k

Bug fixes

effectcheck 0.3.0j

Bug fixes

effectcheck 0.3.0i

Bug fixes and cleanup

effectcheck 0.3.0h

Bug fixes and cleanup

effectcheck 0.3.0g

Bug fixes and new features

effectcheck 0.3.0f-fix

Bug fixes

effectcheck 0.3.0f

Parser fixes and artifact detection

Addresses 13 false positive ERRORs from downstream v0.3.0c validation (132,537 results, 24 ERRORs). Expected: 24 -> ~10 ERRORs.

Bug fixes

New features

Tests


effectcheck 0.2.8

Design ambiguity improvements

Addresses 399 remaining ERRORs from downstream v0.2.7 audit (132,499 results). Philosophy: compute ALL plausible alternatives under different design assumptions; if ANY alternative matches, downgrade severity.

New features

Bug fixes

Internal

effectcheck 0.2.7

Bug fixes and API improvements

Bug fixes

Documentation


effectcheck 0.2.6

Design ambiguity + decision error fixes

Based on downstream analysis of 132,499 results from 8,415 articles. These changes reduce the ERROR false positive rate from ~3.9% to ~0.8%.

Design-ambiguous t-test downgrade (check.R)

Decision error requires reported p-value (check.R)

r-test global N guard (check.R)

API changes

effectcheck 0.2.5

PDF extraction quality improvements

Based on downstream extraction analysis of 121,040 results from 8,415 PDFs across 7 journals. These changes reduce PDF extraction artifacts affecting statistical parsing from ~6.5% to ~0.6%.

Header/footer stripping (utils-pdf.R)

Dropped decimal recovery (parse.R)

General line-break joining (parse.R)

Standalone page number removal (parse.R)

Computation-guided decimal recovery (check.R — Phase 5B)

New columns

Tests

effectcheck 0.2.4

Validation-driven improvements

Based on comprehensive validation of 19,690 results across 7 journals (downstream).

Bug fixes (Category A — 673 results)

Extraction guards (Category B — 41 PDF extraction artifacts)

New features

effectcheck 0.2.3

New features

Bug fixes

API changes

effectcheck 0.2.2

Bug fixes

New features

effectcheck 0.2.1

Bug fixes

Improvements

effectcheck 0.2.0

New features

Bug fixes

Parser improvements

effectcheck 0.1.0