Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
84 commits
Select commit Hold shift + click to select a range
1bac51f
feat(literature): add snowball subcommand (backward/forward/citances)
cjer Jul 1, 2026
e71ad8d
feat(corpus-research): corpus-research skill — methodology + generic …
cjer Jul 2, 2026
ef47fbc
feat(corpus-research): A2 clean-run fixes — validate.py merge-gate, a…
cjer Jul 3, 2026
7bc49f0
feat(corpus-research): coverage-signal tuning from A2 clean-run eval …
cjer Jul 3, 2026
51f07a2
feat(corpus-research): run-mechanics fixes — canon-anchor REQUIRED + …
cjer Jul 3, 2026
f22849f
feat(corpus-research): run README dashboard requirement — ≤25-line st…
cjer Jul 3, 2026
31d72ff
fix(corpus-research): wire run dashboard into the harness — Step 0 wr…
cjer Jul 3, 2026
bc3e773
fix(corpus-research): resolve_titles separates ERRORED from UNRESOLVE…
cjer Jul 3, 2026
642e4d7
feat(corpus-research): acquisition doctrine — match modalities to whe…
cjer Jul 5, 2026
af64813
feat(corpus-research): index-independence rule — the index does not d…
cjer Jul 5, 2026
e0490c0
feat(corpus-research): work-identity model — S2 corpusId as primary k…
cjer Jul 5, 2026
ea95744
feat(corpus-research): incremental-write worker discipline — append-a…
cjer Jul 5, 2026
a58d2b4
feat(corpus-research): post-pair improvement pass (batch 1)
cjer Jul 6, 2026
ebb59a4
feat(corpus-research): cost gate + calibration corrections (Round A e…
cjer Jul 6, 2026
f074243
feat(literature): consume paper-finder enrichment — origins, rejected…
cjer Jul 6, 2026
e25c62c
feat(corpus-research): skill batch 3 — calibration-backed doctrine
cjer Jul 6, 2026
cb4f2fb
docs(corpus-research): calibrated CR pairing rule (de-correlation re-…
cjer Jul 6, 2026
f424188
docs(corpus-research): make the two implicit vocabularies explicit
cjer Jul 8, 2026
80a14f7
fix(review): harden enriched-server consumer + genericize thread-shap…
cjer Jul 8, 2026
c3adc07
fix(corpus-research): cost gate numbers dedup-corrected (~$230-330/ru…
cjer Jul 8, 2026
275c11a
fix(corpus-research): sweeps request include-rejected SAMPLE not summary
cjer Jul 8, 2026
e1bfd7e
fix(corpus-research): sweep at modest concurrency (~3-4) — backend co…
cjer Jul 8, 2026
a610d25
feat(batch-4/A): staged sweep policy + query floors + anchor diversit…
cjer Jul 9, 2026
820defb
feat(batch-4/B): verdict anchoring rule + required truncation term + …
cjer Jul 9, 2026
7b53e55
feat(batch-4/C2+C4): re-read phase reference at the phase (prose deca…
cjer Jul 9, 2026
9d62f96
feat(batch-4): librarian class — curation classifies thread-surveys/r…
cjer Jul 9, 2026
b57c37e
fix(batch-4): librarian sub-classes + anchor axes are EXEMPLARS not e…
cjer Jul 9, 2026
cf65510
batch-4/F tier-1: cut dead code — knowledge.py query surface (1/130 m…
cjer Jul 10, 2026
a6ea359
batch-4/F tiers 2-3: doctrine dedup (one authoritative copy + pointer…
cjer Jul 10, 2026
ce2529b
batch-4/F items 17+18: evidence_depth vocabulary unified (tier = rele…
cjer Jul 10, 2026
73e255f
batch-4/F items 21+22: playbook §6 compressed (one head/tail+anchor r…
cjer Jul 10, 2026
e979e66
feat(batch-4 scripts): knowledge.py view+grep (E1), fulltext ACL/OA r…
cjer Jul 10, 2026
c4fb8fc
fix(triggers): non-paper corpus subjects — 'map the landscape of <too…
cjer Jul 10, 2026
f3dbe4e
fix: description under 1024 chars (validator caught the trigger addit…
cjer Jul 10, 2026
74502dc
fix(report_gate): boundary regex accepts valid reader-facing phrasing…
cjer Jul 10, 2026
4971bfb
batch-5 A: vaults become a skill primitive (references/vault.md + scr…
cjer Jul 11, 2026
b8fb5f4
vault boundary = rounds-of-work, not sessions (user call): same-sessi…
cjer Jul 11, 2026
75198b4
batch-5 B+D: coverage doctrine (enumerator diversity 1/psi, denominat…
cjer Jul 11, 2026
6b78f65
batch-5 C: skill-ecosystem audit edits — S2 scale boundary in semanti…
cjer Jul 11, 2026
98cd050
batch-6 head start: DELIVERY cluster + amendment semantics + wording …
cjer Jul 11, 2026
2b6a291
batch-6: DISPUTED resolution overlay + living-report layering doctrine
cjer Jul 11, 2026
5932382
fulltext.py: fetch+cache PDFs as RAW BYTES (utf-8 decode-replace corr…
cjer Jul 11, 2026
56d6e9c
coverage-playbook B1 promoted: psi measured on known-denominator gold…
cjer Jul 11, 2026
721b6fe
playbook: mechanism-diverse modality = measuring instrument (psi cons…
cjer Jul 11, 2026
808c993
psi claim CORRECTED after stress test (user skepticism vindicated): d…
cjer Jul 11, 2026
c5c960d
playbook: least-biased-anchor wording; B3 our-data receipt
cjer Jul 11, 2026
cff6441
relevance.py: inject title [T]-side (two rounds' fleets dropped it de…
cjer Jul 12, 2026
63cf168
psi replication (oss, 2nd population): pair choice swings -28%..+64% …
cjer Jul 12, 2026
46aa6cd
playbook: per-cell ledger stop rule (UPR part 2 receipts: sim + round…
cjer Jul 12, 2026
f04897b
vault contract: answer-invalidation step for corpus-changing rounds (…
cjer Jul 12, 2026
d6fca46
playbook: persona-shifted canon check (measured anchor move; found a …
cjer Jul 12, 2026
b6c2214
disagreement axes: candidate generation from adjacent ring + review r…
cjer Jul 12, 2026
6fc04e8
reviews.py: OpenReview review-register modality (validated 8/8 vs han…
cjer Jul 12, 2026
5376e1d
reviews.py improvement round: all-candidates v1 resolution (version m…
cjer Jul 12, 2026
9d4f976
reviews.py rung 0: deterministic corpusId->S2 DBLP->forum resolution …
cjer Jul 12, 2026
675a9c8
reviews.py: stream rows + progress per paper (batch death keeps compl…
cjer Jul 12, 2026
202ab5d
coverage playbook: EXPECTATIONS preamble (E1-E4 decisions, plain stat…
cjer Jul 12, 2026
d45f9be
playbook: define 'regime' explicitly (user-caught undefined internal …
cjer Jul 12, 2026
2f057ac
vault.md: eval boundary stated precisely (own-gold never; cross-threa…
cjer Jul 12, 2026
c54c0ad
vault.py verify: derived-layer staleness check (registry row-counts +…
cjer Jul 12, 2026
ae82d8e
vault.md: charter provenance explicit (per-round versioning; inherite…
cjer Jul 12, 2026
03622f4
batch-7 C2: vault.db — disposable sqlite derived index (rows/question…
cjer Jul 13, 2026
4910f38
batch-7 A3/A4/A6: rebuild prints its own delta; canonical fleet-waite…
cjer Jul 13, 2026
d97c45f
batch-7 A1/A2/A5: sweep.py (canonical parallel sweep driver w/ escala…
cjer Jul 13, 2026
b95b2c3
batch-7 B2/C3/C5/C6: contract-as-code fold gate (charter+as_of requir…
cjer Jul 13, 2026
6638e15
batch-7 B1/C4: verdict_gate (5 default-promise checks; discriminated …
cjer Jul 13, 2026
f6283b3
Step-0 guard: drift-prone boundaries always put to the user at the sc…
cjer Jul 13, 2026
72c1dc2
Fold gate accepts charter_inherited_from as charter provenance
cjer Jul 13, 2026
d7181e5
Tool discoverability: scripts index + name tools at their decision po…
cjer Jul 13, 2026
c393a5c
reviews.py coverage map (hosted != readable) + scripts-index boundary…
cjer Jul 14, 2026
a1669f8
Batch-9 doctrine: trust-boundary gates + informed defaults (all recei…
cjer Jul 15, 2026
f0ebf8f
Batch-9 §3: inspect.py toolkit (peek/tally/check) — OPTIONAL, cross-d…
cjer Jul 15, 2026
5eba777
Fix B8-5: rename inspect.py -> inspect_data.py (stdlib shadow that br…
cjer Jul 15, 2026
146f232
Playbook: gap/future-work answers must go back to the papers' own ope…
cjer Jul 15, 2026
031eb9e
Reframe gap-answer line as an INSTANCE of evidence-tier-matching (gen…
cjer Jul 15, 2026
cb4bd7e
coverage-playbook: EVIDENCE-TIER RUBRIC v0 — litmus + per-claim-shape…
cjer Jul 19, 2026
07ea91a
scripts/preflight.sh: portable thread/environment preflight (tooling,…
cjer Jul 20, 2026
f747400
docs: preflight.sh in the scripts index (the consolidated portable co…
cjer Jul 20, 2026
a11fd03
corpus-research: learn-mode/co-work conventions v0 reference page (le…
cjer Jul 22, 2026
43760e8
corpus-research: cost-actual-at-close convention (mandatory close lin…
cjer Jul 22, 2026
01ea638
corpus-research: the EVIDENCE LAYER batch (S-C ruling 2026-07-27) — e…
cjer Jul 27, 2026
d9de52d
corpus-research 12b: the round-14 fix round — query-spec ruling (sear…
cjer Jul 27, 2026
038beaf
corpus-research batch 13: support-gate redesign (quality-not-share fl…
cjer Jul 29, 2026
b874098
README: corpus-research preview entry on this branch (what it is + ma…
cjer Jul 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,13 @@ npx skills add allenai/asta-plugins --all -g
> /plugin install asta-preview
```

- **Corpus Research** (this branch only, preview) - Build and interrogate a curated,
evidence-linked literature corpus over multiple rounds (acquisition, judging, extraction,
reports). Dev setup: `make install`, then launch Claude Code with the plugin loaded
(see DEVELOPER.md for the recommended alias). Before a session, run the environment
check from your working folder:
`bash <repo>/plugins/asta-preview/skills/corpus-research/scripts/preflight.sh feat/corpus-research`
(validates the asta CLI, your `S2_API_KEY`, auth token, and the paper-finder endpoint).
- **Find Literature** - General paper searching and citation finding
- **Literature Report Generation** - Comprehensive report writing with synthesis
- **Run Experiment** - Computational experiments with automated report generation
Expand Down
333 changes: 333 additions & 0 deletions plugins/asta-preview/skills/corpus-research/SKILL.md

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# Learn-mode / co-work conventions (v0)

Everything here is OFFERED, never required (P3): conventions a session composes when the user
is learning through the corpus, not just interrogating it. The frame: co-work between two
agents, each holding (1) their own understanding and (2) a model of the other's — and both can
be stale. Neither side is an oracle. The mode's defining behavior is healthy constructive
doubt in BOTH directions, with repair when the models diverge.

## 1. Learner ledger — `vault/learner.md`
The system's model of the user, made INSPECTABLE AND NEGOTIABLE. Record the user's stated
beliefs, assumptions, and compass anchors VERBATIM (their words, quoted — never paraphrased),
each with date + status `{active | tested | revised | retired}`. Append when the user states a
belief or intrigue; the USER owns edits — a user disagreeing with an entry is itself signal,
not an error. Sections:
- **Compass** — what the user says the work is ultimately FOR; check load-bearing decisions
against it.
- **Expertise map** — user-DECLARED expert/novice areas. Scaffolding that helps novices hurts
experts (expertise-reversal): explanation depth and pushback framing key off this field.
Never infer entries; ask or wait for declaration.
- **Beliefs** — dated verbatim entries with status. A belief gets TESTED only when a question
touches it (no automatic misconception hunting); a status flip cites the evidence.

## 2. Turn-stance read + calibrated pushback
Before executing a user turn, read it: **RULING** (execute) / **HYPOTHESIS** (test it) /
**REACTION** (probe what triggered it). Silent by default — surface the read ONLY when it
changes the action. Push back only when expected value exceeds the interruption cost (the
mixed-initiative when-rule): one sentence, naming the consequence the user may not see
("replacing loses the comparison — still replace?"). Proportionality is the zeroth row: doubt
scales with stakes; never interrogate 1+1. **Misfire log**: record every pushback + outcome
`{upheld | overruled}` in the session notes — wrong pushbacks are data, not embarrassments.

## 3. Premise-stating on action
At load-bearing moves, state the premise: "acting on X because I believe Y" — the
instruction-side twin of the how-performed note. It hands the user a hook to catch the
session's stale model before damage.

## 4. Trajectory links
QUESTIONS.log entries gain one optional field: `spawned_from: <id of parent question/turn>`.
Morphing questions stay one traceable arc instead of disconnected asks. One field, no new file.

## 5. Two-way uncertainty (seed)
State the session's own confidence where the probe is licensed — membership, synthesis,
disagreement shapes — and NEVER absence-like shapes (measured: confident-but-wrong on every
absence flip). Ask for the user's uncertainty SPARINGLY, at genuine junctures — uncertainty is
a stage signal to steer by, not a form to fill.

## 6. Gricean norm block (adopted from measured results: +27% task accuracy, +8% appropriate
## clarification when stated explicitly)
Include in learn-mode session guidance:
> Say as much as the exchange needs and no more. Say only what you have evidence for — and
> mark what you don't. Make it relevant to what the user is actually pursuing (their words,
> their compass). Be orderly and plain. When an instruction is ambiguous and the ambiguity
> changes the outcome, ask ONE targeted clarifying question instead of guessing.

## 7. SRL-lite session frame
Open with a 2-line goal in the user's words ("what do you want to understand by the end?").
Close with a 2-line reflection (what moved, what's still open) — appended to the ledger arc,
not a ceremony.

## 8. The erosion warning (design principle, not a primitive)
Well-powered RCTs: assistance can raise in-task performance while ERODING persistence,
unassisted skill, and metacognitive calibration. At learning junctures prefer prompting the
user's own reasoning (a self-explanation one-liner: "what would X predict here?") over handing
the answer. Corrective friction is a feature.
69 changes: 69 additions & 0 deletions plugins/asta-preview/skills/corpus-research/references/codebook.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
# Codebook — grounded aspect vocabulary, DERIVED from this corpus

When a question needs content families (which abilities/methods/phenomena/techniques a corpus
covers), you need a shared tag vocabulary. DERIVE it from THIS corpus — never import a codebook
from another thread (that's a train-test leak and it won't fit).

## Derivation (grounded coding)
1. **Open-code** a sample (titles+abstracts of the relevant set): note the phenomenon/method each
analyzes, in the papers' own terms.
2. **Cluster** the open codes into ~10–20 families; each family = {name, one-line definition,
a few exemplar corpusIds}.
3. **Name + freeze** as `codebook@v1`; write `codebook_version` into thread.json.
4. **Apply by exact-match** to every relevant paper as an ordered tag batch
(`<run>/tag-batches/<NN>-<name>.jsonl`: {corpusId, primary_family, secondary_families, confidence}).
5. **Monitor "Other"** — genuine "other/unclassifiable" (excluding empty-abstract) growing past
~10% of the relevant set triggers re-derivation: re-open-code the "Other" pile → codebook@v(N+1)
→ re-tag → rebuild. Bump the version; the substrate stamps it on every record.

## Two failure modes the gates catch (both real)
- **Untagged ≠ Other.** When the relevant set EXPANDS (new acquisition rounds), the new papers
must be tagged before any distribution analysis, or an inflated "Other" bucket silently distorts
the family distribution. `substrate.py`'s TAG_GATE fails <90% real-family — do not quote a
distribution until it passes.
- **Over-conservative tagging.** A batch that dumps classifiable papers into "Other" fails the
same gate. Re-tag them (they usually fit existing families) rather than inventing families.

## Parametric family anchor — REQUIRED gate after codebook@v1 (§anchor)
A corpus-derived codebook is circular by construction: a whole missing family produces no
"Other" growth, no tag-gate failure, nothing. The anchor is the validated check for that class
(caught 4 real families hidden in a real run's "Other"; correct null on a good codebook;
charter-conditioning is the false-positive control — 15/20 spurious naive → 0 conditioned):
> **Protocol.** Reading ONLY the thread charter (question, deliverables, out_of_scope_families)
> — NOT the codebook or corpus — enumerate 12–25 content families you'd expect this literature
> to span, with one-line definitions; WRITE THEM DOWN before opening the codebook. Then align
> each to the codebook: PRESENT / SUBSUMED / ABSENT. For every ABSENT family, title-scan the
> run's candidates and core: (a) papers in-core → carve out the family or add it as a NAMED
> residual (never silently in Other); (b) papers in candidates but not core → audit the
> curation/scope decision; (c) papers nowhere → thin-literature note or an acquisition probe.
> For SUBSUMED families with substantial in-core mass, consider promotion at codebook@v2.
> Record the table in the run's coverage evidence; the gate FAILS if any ABSENT-case-(a) family
> is left uncarved and unnamed.
The anchor is one-directional (it misses families you didn't think to enumerate) — it
complements grounded open-coding, never replaces it.

## A codebook is also a scope audit
Deriving families exposes clusters that don't belong to the thread's phenomenon (an orthogonal
subfield, a methodology-only cluster). Those become `scope.out_of_scope_families` in thread.json
(when scope.axis == "separate") — the family-based half of the scope test.

## Methods/entities get their own codebook
The same recipe derives a METHODS codebook (normalize free-text method strings into countable
families) or an ENTITY vocabulary — whatever the question aggregates over. Derive per axis.

## Canonicalization maps are DATA — gate them like data
Entity/name canonicalization (model names, method names, dataset names) uses a **2-level scheme**:
raw string → `{canonical_name, family}` (spelling variants → one canonical; versions → one
family). Rules, each learned from a real failure:
- **Frozen artifact, never inline.** The map is a versioned FILE (`canon-map.json`) applied
deterministically [T]. Inline normalization ships its errors silently; a file is a 30-second
audit.
- **Attested names only.** `canonical_name` must appear (modulo punctuation/case) somewhere in
THIS corpus — a title or an extraction. NEVER invent a name; when none is attested, keep the
raw string as the canonical and assign only `family`. Real failure: parameter sizes became
phantom versions ("FooCoder 33B" → "Foo 33", "a 1.1B Bar-architecture model" → "Bar 1.1").
- **Machine self-check → review queue.** Scan the map for unattested canonical names (normalized
substring check against the corpus vocabulary); fix or revert each hit; surface UNCERTAIN
mappings to the user at a beat — a review queue, not silent auto-accept.
- **General rule: pair every "produce X" with "gate X".** Any derived mapping/table ships only
after an acceptance check on the artifact itself — same discipline as merges and substrates.
Loading