Skip to content

adr: ranked search over alternative labels, and duplicate-hit disambiguation #92

Description

@maehr

Context

ADR-0007 gave Work one flat alternative_labels list and deferred two questions to the
moment the registry gained ranked search. Its own words: option 4 was "rejected on authoring
cost, not on merit", and open question 2 says "Revisit if the browser gains ranked search".

/find/ is that moment. It ranks works on an explicit seven-tier ladder, so the deferred
questions now have a concrete consumer rather than a hypothetical one.

Two things surfaced while building it.

An abbreviation and a title behave differently under ranking. NE should match the
Nicomachean Ethics exactly and nothing else. Nikomachische Ethik should also match on a
prefix or a token. One flat list cannot express the difference, so the finder treats every
entry the same way.

The cost is observable today. Because no work carries alternative_labels yet, the query
NE 1094a1 resolves to the New Testament: ne is a literal prefix of new testament,
which is tier 5, and nothing better matches. Seeding NE as an alternative label fixes this
case, because an exact alternative label is tier 3 and outranks a prefix. But the shape of
the failure is general: a two-letter abbreviation is a prefix of many titles, and a flat
list gives the ranker no way to say "this entry is an abbreviation, match it exactly".

A duplicate hit still has nothing to disambiguate it. ADR-0007 allows two works to share
an alternative label and requires consumers to treat the match as ambiguous. /find/
does: it returns both. But the rows carry only the label, the creator, the status, and the
key. ADR-0007 listed this as an open follow-up and did not solve it.

A related asymmetry appeared. When one work's preferred label equals another work's
alternative label, the ladder decides rather than asks, because tier 2 beats tier 3. If
Spinoza's Ethics is registered, Ethics returns Spinoza alone and never mentions
Aristotle. That follows the ranking rules, and it may be wrong for a reader. ADR-0007's
homonymy rule speaks only about two alternative labels.

Options considered

  1. Do nothing. Keep one flat list. The finder keeps ranking every entry the same way.
    Cheapest, and defensible while the registry holds twelve works. It does not scale to the
    fields the roadmap targets, where abbreviations are the common form of citation.
  2. Split the field into alternative_labels and abbreviations. This is ADR-0007's
    rejected option 4, now with a consumer that would use the distinction. It doubles the
    authoring decision, and ADR-0007's objection stands: LXX is an abbreviation,
    Septuaginta is a title, and Sept. is arguable.
  3. Keep one list and derive the distinction. Treat a short, mostly-uppercase entry as an
    abbreviation and require an exact match for it. No authoring cost, no schema change. The
    heuristic is the whole risk, and it would be a rule the standard does not state.
  4. Language-tagged label objects, ADR-0007's open question 1. It answers a different
    question, but it touches the same field, and doing both at once avoids two breaking
    changes to one shape.

Recommendation

Decide options 2 and 3 together, and decide the duplicate-hit presentation with them.

Option 3 is worth serious weight. It costs no schema change and no authoring burden, and the
rule can live in the ranker rather than in the record. Its weakness is that a heuristic in
one client is not a contract; if the distinction matters, it belongs in the data.

Whichever is chosen, the duplicate-hit question needs an answer of its own: a shared label
should carry enough context to choose. The creator is usually enough, and the finder already
shows it. The gap is a work with no creator.

Do not decide language tagging here. It is a larger question and ADR-0007 already frames it.

Expected consequences

  • The ranking ladder gets a stated contract instead of an implementation detail.
  • If the field splits, existing records stay valid, because both fields are optional. The
    authoring guidance in get-started/authoring.md needs a boundary rule.
  • If the distinction stays a heuristic, it must be written down as non-normative, so a second
    client is not obliged to copy it.
  • Identity is untouched either way. ADR-0002 fixes the reference UUID seed, and no label
    reaches it.

Notes

/find/ merged in #93, so the consumer this decision needed now exists and can be pointed
at: rankWorks in src/lib/find.ts implements the seven-tier ladder, and
src/lib/find.test.ts pins the homonymy and preferred-versus-alternative cases described
above.

Nothing is blocked. The finder is correct under either outcome; only the ranking contract
changes.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    adrArchitecture Decision RecordenhancementNew feature or requestpost-v0.1.0Deferred past the v0.1.0 baseline. Revisit if the need arises.standardThe published specification and schemas

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions