---
title: "Should the vault's single source_tier field split into two axes — source reliability separate from claim credibility — the way intelligence doctrine and RAG both do?"
type: "question"
status: "open"
date_raised: "2026-07-11T00:00:00.000Z"
tags: ["vault-design","source-tiers","provenance","epistemics","meta"]
---


The Admiralty-Code hop ([[observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model]]) surfaces a direct challenge to the vault's own design. Both Cold-War intelligence doctrine ([[claim-admiralty-code-grades-sources-on-two-independent-axes]]) and 2025 retrieval research ([[claim-reliability-aware-rag-estimates-source-reliability-separately-from-relevance]]) grade a report on **two** axes — how reliable the *source* is, separately from how credible the *specific claim* is. The vault's `source_tier` fuses these into one number.

**The question.** Should a claim-note carry two fields — e.g. a source-reliability tier and a per-claim credibility/corroboration grade — instead of one `source_tier`?

**The case for.** The vault already half-does this: `audit_status`, `[unverified-quant]` flags, and the "one confirmed source vs uncorroborated" distinction are credibility-of-the-claim signals living *outside* `source_tier`. Making the second axis explicit could sharpen retrieval gating.

**The case against.** [[claim-source-reliability-and-credibility-are-not-judged-independently|Trained analysts empirically cannot keep the two axes independent]] — they over-weight source track record and avoid inconsistent pairings. If humans (and, per [[claim-document-text-degrades-llm-source-authority-judgment|AuthorityBench]], models) can't hold the axes apart, a second field may add false precision rather than honesty. The single tier may be an *admission* that the split doesn't survive contact with a real evaluator.

**What I'd need to answer it.** A pass over how `source_tier` + `audit_status` + flags actually interact in ~20 notes; whether the fusion has ever caused a mis-grade; and whether the retrieval layer (spec §7's gating rule) would benefit from a separate credibility axis. This is a design decision for Cali, not a unilateral schema change.

**Candidate next move.** Draft a one-page design memo weighing the two-axis option against the current fused tier, citing the three underlying claim-notes; leave the schema untouched until Cali rules.

## Progress

**2026-07-21 (partial — stays open).** A batch capture
(`10-inbox/raw/2026-07-20-should-the-vaults-single-source-tier-field-split.md`)
added a third and fourth independently-converging domain and a check on what
the existing "axes leak" finding actually recommends: evidence-based
medicine's GRADE framework
([[claim-grade-splits-quality-of-evidence-from-strength-of-recommendation]])
and the CRAAP library-science test
([[claim-craap-test-splits-authority-from-accuracy]]) both draw the same
source-vs-content line independently; and Kelly et al.'s own response to
"raters can't keep the axes apart" turns out to be a richer nine-cell matrix,
not a call to collapse
([[claim-kelly-et-al-propose-richer-joint-matrix-not-axis-collapse]]). A
search for practitioner opinion on the merge-vs-split question found writers
arguing only to keep the axes separate
([[claim-no-practitioner-source-found-advocating-merging-reliability-credibility-axes]]),
though that search was not exhaustive.

This strengthens the case-for column but does **not** close the question: no
inter-rater-reliability, cognitive-load, or rater-burden data on two-axis
systems turned up in this pass either (checked specifically in Kelly et al.
and came up empty), so the cost side of the tradeoff — whether a second field
would sharpen retrieval gating or just add a number nobody keeps honestly
independent — remains as unaddressed as it was on 2026-07-11. Still a design
decision for Cali, not something four converging citations settle by
themselves. Left open.

**2026-08-07 (partial — stays open, new third option surfaced).** Reading
Benjamin Icard's own 2024 primary
([[claim-icard-2024-dynamic-logic-makes-credibility-primary-reliability-secondary]])
shows the "Icard (2023, 2024)" citation Kelly et al. point to is not the
independence-preserving richer-matrix design the vault's existing note
characterized it as (that characterization is accurate to what Kelly et al.
say, just one citation-hop short of what Icard actually built). Icard's own
proposal is a third shape neither column above considered: two named
quantities, formally *not* independent by construction, with reliability
demoted to a dynamic update operator on credibility rather than a coequal
second axis. This doesn't answer the design question — whether the vault's
`source_tier` should fuse, split, or asymmetrically-update is still Cali's
call — but it means "split into two independent axes" was never the only
alternative to fusion, and any design memo on this question should now weigh
Icard's asymmetric-update shape alongside fuse/split. Left open.

**2026-08-07 (partial — stays open, the leak now has a century-old name).**
A hop capture landing Thorndike's 1920 primary
([[claim-thorndike-1920-halo-effect-ratings-too-high-and-too-even]],
[[observation-halo-effect-names-the-2025-reliability-credibility-leak]])
identifies the "raters can't keep the axes apart" evidence in the case-against
column as an instance of the **halo effect** — a robust, cross-domain,
century-old psychometric bias, not a quirk of intelligence analysts. This
strengthens the case-against a two-independent-axes split: the axes leak not
because a particular evaluator population is careless but because a global
impression contaminating independent trait ratings is a general and durable
feature of human judgment. It does not settle the design question (fuse vs.
split vs. Icard's asymmetric update remains Cali's call), and it says nothing
about the still-unaddressed cost side (rater burden, whether a second field
sharpens retrieval gating). Any design memo should now note that the leak it
must design around is the halo effect by name. Left open.

**2026-08-15 (partial — stays open, a third design dimension surfaced).**
Reading Samet (1975) at the primary (promotion of the 2026-08-11
channel-capacity hop) adds a dimension neither the fuse/split nor the Icard
asymmetric-update framing considered: **granularity per axis**. The Admiralty
Code's real in-house critique was not "too many axes" but "too few rungs" —
Samet argued the scales were too *coarse* and should be finer, not fused
([[claim-samet-1975-argued-for-more-rating-categories-not-fewer]]), grounding
the argument in Bendig (1954) and the information-theoretic ceiling
[[claim-miller-1956-channel-capacity-limits-absolute-judgment-categories|Miller (1956) named as channel capacity]]
(lineage: [[observation-admiralty-code-scale-length-descends-from-millers-channel-capacity]]).
Any design memo on this question should now weigh *scale length* — how many
levels each field carries — alongside fuse-vs-split-vs-asymmetric-update. This
does not settle the design question (still Cali's call) and says nothing about
the still-unaddressed cost side. Left open.

**2026-08-18 (partial — stays open, correcting the entry immediately above and
adding one more data point on the "how many axes" question).** The 2026-08-15
entry's summary of Samet — "argued the scales were too coarse and should be
finer, not fused" — is stale relative to a correction already applied to
[[claim-samet-1975-argued-for-more-rating-categories-not-fewer]] on 2026-08-16:
Samet's own implications section (p. 20) explicitly recommends fusion as well
as finer resolution — "the two-dimensional evaluation should be replaced" by a
single quantitative rating, with his subjects voting 21–16 in favor. So Samet
is not a clean example of "finer, not fused"; he is a Tier-1 practitioner
source arguing for *both at once* — one axis, made finer. A same-day bridge-check
([[observation-kelly-samet-cosine-pairing-real-link-opposite-axis-prescription]])
adds that Kelly et al. (2025) cite Samet directly in their own §2 for the
axes-leak evidence, yet their discussion's own forward pointer (Icard's richer
joint matrix) runs the opposite structural direction — more explicit joint
categories across two axes, not one fused axis. So the "add structure" and
"fuse the axes" columns above both now have their strongest respective
advocates reading the same underlying evidence and prescribing oppositely,
inside literatures that do not engage each other's specific fix. This narrows
what a design memo can claim ("the field agrees more granularity helps" is not
supportable — granularity-of-what remains contested) without resolving fuse vs.
split vs. asymmetric-update, which is still Cali's call. Left open.

---

> [!warning] Correction appended 2026-09-13 (propagation-repair)
> Propagating the CORRECTED 2026-07-22 cross-model audit (auditor claude-fable-5,
> writer claude-sonnet-5) of
> [[claim-kelly-et-al-propose-richer-joint-matrix-not-axis-collapse]]. The
> progress entries above stand character-for-character — this file records the
> answer as it was given — but the **2026-07-21** entry's characterization of the
> nine-cell matrix is superseded.
>
> **Was:** "Kelly et al.'s own response to 'raters can't keep the axes apart'
> turns out to be a richer nine-cell matrix, not a call to collapse" — the
> matrix framed as Kelly et al.'s own prescription, answering their own finding.
>
> **Now:** two separate attributions in that sentence were too strong. (1) The
> 3 × 3 matrix is **Icard's** proposal (Icard 2023, 2024), which Kelly et al.
> merely *point to* in their general discussion as a comparison target for
> future work — "future studies could compare the reliability and perceived
> usefulness of the Admiralty Code to alternative methods that encode
> qualitative meaning at the 'cell' level" — not their own prescription. The
> claim note's verb was softened from "propose" to "point to" accordingly.
> (2) The "raters can't keep the axes apart" evidence is **prior work Kelly et
> al. review in their §2** (Baker et al. 1968; Miron et al. 1978; Samet 1975;
> Mandel et al. 2023), not their own experimental finding; their experiments
> concern decoding, adding that trustworthiness ratings over-weight source
> reliability. So it is not "their own response" to "their own" finding on
> either side. The same audit also fixed the quotation, which was not verbatim
> despite the capture's claim, against the publisher HTML and the PDF
> (doi:10.1017/jdm.2025.10007).
>
> **What did not move.** The directional point the 2026-07-21 entry drew on is
> unaffected: the paper's discussion, faced with the axes-leak evidence, still
> runs toward more explicit joint granularity rather than a single collapsed
> score. So that entry's conclusion — strengthens the case-for column, does not
> close the question — stands, as does its note that no inter-rater-reliability
> or rater-burden data turned up. The claim note's filename is unchanged, so the
> link above still resolves. The **2026-08-07** and **2026-08-18** entries
> already use the corrected voice ("the 'Icard (2023, 2024)' citation Kelly et
> al. point to", "their discussion's own forward pointer") and need no
> annotation; the 2026-08-07 entry's separate finding — that Icard's *own*
> primary builds an asymmetric-update shape rather than two independent axes —
> is a different, later correction and is untouched here.
>
> **No routing change.** `status: open` preserved: this correction adjusts whose
> prescription the nine cells are, not whether the vault's `source_tier` should
> fuse, split, or asymmetrically update. Still Cali's call, and the cost side
> remains unaddressed.
