Persist relevance scores in the SQLite semantic ref index - #324
Bernhard Merkle (bmerkle) wants to merge 1 commit into
Conversation
SqliteTermToSemanticRefIndex dropped the score of a ScoredSemanticRefOrdinal on write and returned a fabricated 1.0 on read, so rankings diverged from the memory backend. Add a score column (migrating existing databases) and persist/return it in add_term, add_terms_batch, lookup_term, serialize and deserialize. Fixes microsoft#321 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The schema migration can fail when multiple providers concurrently initialize the same legacy database.
Review effort: Balanced
Findings: 1
Open (2)
What changed in this PR
Persists semantic-reference relevance scores in SQLite to match the memory backend.
Changes:
- Adds and migrates the SQLite score column.
- Preserves scores across writes, lookups, and serialization.
- Adds backend parity and migration tests.
| File | Description |
|---|---|
src/typeagent/storage/sqlite/schema.py |
Adds score schema and legacy migration. |
src/typeagent/storage/sqlite/semrefindex.py |
Persists and returns relevance scores. |
tests/test_semrefindex.py |
Tests score preservation and migration. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| columns = [row[1] for row in cursor.execute("PRAGMA table_info(SemanticRefIndex)")] | ||
| if "score" not in columns: | ||
| cursor.execute( | ||
| "ALTER TABLE SemanticRefIndex ADD COLUMN score REAL NOT NULL DEFAULT 1.0" | ||
| ) |
There was a problem hiding this comment.
unlikely but seems like a simple enough change
| import sqlite3 | ||
|
|
||
| from typeagent.storage.sqlite.schema import init_db_schema |
| CREATE TABLE IF NOT EXISTS SemanticRefIndex ( | ||
| term TEXT NOT NULL, -- lowercased, not-unique/normalized | ||
| semref_id INTEGER NOT NULL, | ||
| score REAL NOT NULL DEFAULT 1.0, |
There was a problem hiding this comment.
Won't defaulting the score to 1.0 might hide more valid results? Should this be 0 or some really small value?
| columns = [row[1] for row in cursor.execute("PRAGMA table_info(SemanticRefIndex)")] | ||
| if "score" not in columns: | ||
| cursor.execute( | ||
| "ALTER TABLE SemanticRefIndex ADD COLUMN score REAL NOT NULL DEFAULT 1.0" | ||
| ) |
There was a problem hiding this comment.
unlikely but seems like a simple enough change
| ) -> tuple[SemanticRefOrdinal, float]: | ||
| if isinstance(ordinal, ScoredSemanticRefOrdinal): | ||
| return ordinal.semantic_ref_ordinal, ordinal.score | ||
| return ordinal, 1.0 |
There was a problem hiding this comment.
Maybe make this a constant and reuse this in/from schema.py when initializing the SemanticRefIndex table.
| # Fallback for direct integer | ||
| semref_id = semref_ordinal_data | ||
| insertion_data.append((term, semref_id)) | ||
| score = 1.0 |


Fixes #321.
SqliteTermToSemanticRefIndexdiscarded the score of aScoredSemanticRefOrdinalon write and returned a hardcoded1.0on read, so result ordering differed from the memory backend.score REAL NOT NULL DEFAULT 1.0toSemanticRefIndex; existing databases are migrated ininit_db_schemaviaALTER TABLE(old rows get 1.0).add_term,add_terms_batch,lookup_term,serialize,deserialize; order lookups byrowidto match the memory backend's insertion order.Note: the "duplicate pairs" divergence mentioned in the issue does not occur — the table has no unique constraint, so
INSERT OR IGNOREnever dedupes and both backends append duplicates. A test pins this.🤖 Generated with Claude Code