diff --git a/docs/stockfit/README.md b/docs/stockfit/README.md index 49628d1..60c3f39 100644 --- a/docs/stockfit/README.md +++ b/docs/stockfit/README.md @@ -94,6 +94,11 @@ try { - [`methodology.md`](methodology.md) defines the signal and time gates. - [`case-study-draft.md`](case-study-draft.md) is an unpublished portfolio - narrative based on the verified live run. + narrative combining the verified live backtest and data-quality audit. +- [`nav-comparison.svg`](../../artifacts/stockfit-live/nav-comparison.svg) is the + primary portfolio visual; it compares the fundamental rule with its matched + equal-weight baseline. +- [`data-quality.svg`](../../artifacts/stockfit-audit-live/data-quality.svg) is + the supporting audit visual across the frozen twelve-company cohort. - [`learning-guide.md`](learning-guide.md) provides a 90-minute code-tracing session for revisiting the implementation in VS Code. diff --git a/docs/stockfit/case-study-draft.md b/docs/stockfit/case-study-draft.md index 551113a..99c1a37 100644 --- a/docs/stockfit/case-study-draft.md +++ b/docs/stockfit/case-study-draft.md @@ -1,82 +1,134 @@ -# StockFit point-in-time fundamental backtest +# StockFit point-in-time research case study -**Draft — not published** +**Portfolio draft — not published** -## Problem +## Overview -I wanted to extend `backtest-lib` with a realistic external-data example while -keeping the integration small enough to understand and test. The chosen task -was to combine adjusted prices with annual revenue and operating-income facts, -then compare a transparent fundamental ranking with an equal-weight portfolio. +I extended `backtest-lib` with a small, testable external-data workflow for +point-in-time fundamental research. It combines adjusted prices with filing-dated +annual revenue and operating-income facts, runs a transparent ranking strategy +against an equal-weight comparator, and audits whether a broader twelve-company +cohort is structurally usable for the same style of research. -## Why point-in-time matters +The result is an engineering case study rather than an investment claim. The +fundamental rule did not outperform its comparator, and the data audit found +material gaps in three companies. Both outcomes are retained because they make +the research more credible and expose the controls a production process would +need. + +## Evidence at a glance + +| Evidence | Verified result | +| --- | --- | +| Live price coverage | 1,434 shared observations, 4 January 2021 to 18 September 2026 | +| Backtest cohort | AAPL, MSFT and COST | +| Fundamental rule | Top two eligible securities by year-on-year revenue growth plus operating margin | +| Comparator | Equal weight over the same cohort, dates and decision schedule | +| Fundamental result | 157.80% total return; 0.67 Sharpe; -32.14% maximum drawdown | +| Comparator result | 166.27% total return; 0.82 Sharpe; -27.95% maximum drawdown | +| Data-quality audit | 12 companies; 9 pass, 0 review and 3 fail | +| Trial requests | 36 sequential requests with no retries | +| Regression evidence | 424 passed, 3 skipped and 12 deselected; Ruff and Pyrefly passed | + +## Why point-in-time handling matters A backtest can look into the future accidentally if it uses a financial value -before the filing date or applies a later amendment to earlier decisions. The -integration therefore gates each statement by its original filing date and -rolls amended facts back to the value available on each historical decision -date. Missing inputs make a security ineligible rather than silently filling a -value. +before its filing date or applies a later amendment to earlier decisions. This +integration gates each statement by its original filing date and rolls amended +facts back to the value available on each historical decision date. Missing +inputs make a security ineligible rather than silently filling a value. + +The default workflow is offline and uses fixtures explicitly labelled as +invented synthetic data. Live calls require both the `--live` flag and a +process-scoped `STOCKFIT_TOKEN`. Raw provider responses, prices, financial +values and credentials are not retained in the repository. ## Architecture -The example has four narrow layers: +The example keeps network access, transformation, portfolio logic and reporting +separate: -1. `client.py` performs the three approved HTTP reads and validates response - status and shape without embedding credentials. +1. `client.py` performs the three approved HTTP reads and validates status and + response shape without embedding credentials. 2. `transforms.py` aligns prices, reconstructs facts as of each date and builds the library's `MarketView` signals. -3. `strategy.py` contains the top-two ranking rule and the matched equal-weight - comparator. +3. `strategy.py` contains the top-two fundamental rule and the matched + equal-weight comparator. 4. `demo.py` runs both portfolios and emits only derived metrics and a NAV comparison chart. +5. `audit.py` and `audit_cli.py` apply predeclared coverage, integrity, + completeness, filing-timing and provenance checks to an expanded cohort. + +## Backtest result + +The verified live run produced 1,425 completed observations for both strategies +from 4 January 2021 through 4 September 2026. Both portfolios started with +1,000,000 nominal units in cash and rebalanced monthly. + +The fundamental portfolio ended at 2,578,038.91, with an annualised return of +18.23%, annualised volatility of 24.19%, a Sharpe ratio of 0.67 and maximum +drawdown of -32.14%. The equal-weight comparator ended at 2,662,720.38, with an +annualised return of 18.91%, annualised volatility of 20.61%, a Sharpe ratio of +0.82 and maximum drawdown of -27.95%. + +![NAV comparison for the fundamental and equal-weight portfolios](../../artifacts/stockfit-live/nav-comparison.svg) + +The comparator finished ahead and had better risk-adjusted metrics. This result +does not support an alpha claim; it demonstrates that the framework can produce +and preserve an unfavourable but informative comparison. + +## Data-quality result + +Before inspecting outcomes, I fixed a twelve-company cohort spanning technology, +banking, energy, healthcare, consumer and industrial businesses. The audit made +one company, price-history and income-statement request per ticker and evaluated +the responses against predeclared thresholds. -The default path uses explicitly labelled invented fixtures. Live calls require -the `--live` flag and a process-scoped `STOCKFIT_TOKEN`. +Nine companies passed and three failed. JPM and BAC lacked sufficient operating +income coverage to form a usable consecutive annual pair. XOM returned no annual +rows in the trial request, resulting in short-history, incomplete-fact and +no-usable-pair findings. These are narrow observations about the returned trial +data, not judgements about the companies or securities. -## Test strategy +![StockFit data-quality audit by company and check](../../artifacts/stockfit-audit-live/data-quality.svg) -The implementation was developed from failing tests. Unit tests cover request -construction, status and schema failures, secret redaction, price alignment, -filing-date gates, multi-source rollback, missing facts, deterministic -ranking and strategy weights. End-to-end tests assert deterministic offline -artefacts, prevent raw fields from entering outputs, and keep live use opt-in. -The complete repository suite passed 322 tests, with 3 skipped and 12 -deselected; Ruff and Pyrefly also passed. +The audit retained only derived metadata. Across the cohort, median revenue and +operating-income completeness were both 100%, while the company-level median +filing lag was 46.5 days. The complete result and fixed thresholds are recorded +in [`data-quality-audit.md`](data-quality-audit.md). -## Demonstration result +## Verification and reproduction -One verified live run on 21 September 2026 used AAPL, MSFT and COST. The common -price coverage contained 1,434 observations from 4 January 2021 to -18 September 2026; both completed backtest series contained 1,425 observations -through 4 September 2026. +Unit tests cover authenticated request construction, status and schema failures, +secret redaction, price alignment, filing-date gates, amendment rollback, +missing facts, deterministic ranking, audit thresholds and artefact safety. +End-to-end tests run without network access or credentials and assert that raw +provider fields do not enter published outputs. -The fundamental portfolio ended at 2,578,038.91 nominal units from 1,000,000, -a total return of 157.80%. Its annualised return was 18.23%, annualised volatility -24.19%, Sharpe ratio 0.67, maximum drawdown -32.14%, and average turnover 0.37% -per engine period. +At this checkpoint, the repository suite reported 424 passed, 3 skipped and 12 +deselected tests. Ruff formatting and linting, Pyrefly type checking, documentation +builds and cross-platform wheel builds also passed in CI. -The equal-weight comparator ended at 2,662,720.38 nominal units, a total return -of 166.27%. -Its annualised return was 18.91%, annualised volatility 20.61%, Sharpe ratio -0.82, maximum drawdown -27.95%, and average turnover 0.45% per engine period. -In this narrow demonstration, the comparator finished ahead and had better -risk-adjusted metrics. +The offline demonstration and audit remain reproducible from documented commands +in [`README.md`](README.md). Live execution is deliberately opt-in. ## Limitations -This is an engineering demonstration, not evidence of a profitable investment -strategy or production readiness. It uses a fixed three-stock current-ticker -cohort, so survivorship and selection bias are substantial. It does not model -costs, spread, slippage, liquidity, taxes or market impact. Provider history may -be clamped, adjusted-price history may use a current adjustment snapshot, and -restatement reconstruction depends on the source history exposed by the API. -There was no out-of-sample study, universe reconstruction or parameter search. - -## What I learned - -The most important lesson was that data timing is part of strategy logic, not -just data cleaning. A small client and pure transform functions made it possible -to test awkward cases—especially restatements and missing facts—without live API -calls. A deliberately weak but matched comparator also made the result more -honest: the new signal did not beat simple equal weighting in this sample. +- The current-ticker cohorts introduce survivorship and selection bias. +- Three stocks are too few to evaluate a general investment strategy, while + twelve companies are too few to certify provider-wide data quality. +- The backtest omits costs, spread, slippage, liquidity, taxes and market impact. +- Provider history may be clamped, and adjusted-price history may use a current + adjustment snapshot. +- One uniform fundamental fact rule is deliberately conservative across sectors, + particularly banking and energy. +- There was no out-of-sample study, historical universe reconstruction or + parameter search. + +## Engineering takeaway + +Data timing and data quality are part of strategy logic, not preprocessing +details. Narrow interfaces and pure transformation functions made awkward cases +such as restatements, missing facts and incomplete sector coverage testable +without live API calls. A matched comparator and predeclared audit thresholds +kept the evidence interpretable even when the outcomes were unfavourable.