Skip to content

Держать терминальную сессию в README транскриптом, а не картинкой - #14

Merged
DenisDrobyshev merged 2 commits into
masterfrom
check/readme-session
Sep 20, 2026
Merged

DenisDrobyshev merged 2 commits into
masterfrom
check/readme-session

Conversation

@DenisDrobyshev

@DenisDrobyshev DenisDrobyshev commented Sep 20, 2026 •

Copy link
Copy Markdown
Member

Первая страница открывается прогоном manage.py diff — agreement, переходы, пара accuracy, fixed против broken, вердикт с p-значением. Это первое, что читают, и это аргумент за весь инструмент. Не проверялось ничего.

Три способа протухнуть, все молчаливые.

1. Цифру правят, остальные перестают из неё следовать

500 строк при agreement 92.0% — это 40 изменившихся. 40 изменившихся — это 22 + 18 переходов и 29 fixed + 11 broken. 0.7700 → 0.8060 на 500 строках — это 18 строк, то есть 29 − 11.

Теперь это всё утверждается, включая p-значение вердикта: точный двусторонний МакНемар по дискордантным парам даёт 0.006427 против напечатанных p=0.006.

Образец, противоречащий сам себе, — плохая реклама инструменту, чья идея в том, что агрегаты прячут то, что сдвинулось.

2. Два README расходятся

Это транскрипт команды, одинаковый на обоих языках. Перевести его — значит выдумать. Сравниваются символ в символ.

3. Программа перестаёт это печатать

Переименовали метку, расширили колонку — и первая страница показывает вывод, которого больше нет.

Существующие тесты diff утверждают, что agreement и fixed где-то встречаются. Это переживает обе поломки. Поэтому новый тест гоняет настоящую команду на двух настоящих обученных версиях и сверяет форму: каждая метка из образца должна печататься, её значение — начинаться в той же колонке, а блок по разметке — по-прежнему говорить row(s) wrong in, а не case(s) failing in, которое использует ветка evals.

Ничего не было сломано

Все цифры уже сходились, все строки уже совпадали. Это стоит сказать прямо: PR прибивает страницу, которая была права, а не чинит неправую.

Проверено

Сломал копию пятью способами:

что сломано что сказал
цифра в образце правлена agreement and changed disagree
русский README отстал транскрипты разошлись
p-значение подправлено verdict says p=0.001, McNemar gives p=0.006427
метка переименована в diff.py the program no longer prints a 'agreement' line
колонка расширена в diff.py 'changed' puts its value at column 21, the README shows column 17

Разделено по фикстурам: два статических теста не требуют ничего, кроме файлов, и идут полсекунды; живой стоит рядом с TestDiff, где уже есть машинерия обучения модели.

ruff check mlango tests и ruff format --check — чисто. 36 тестов в затронутой области проходят.


Первый прогон был красный, и это моя ошибка

Упало на всех тестовых job'ах:

AssertionError: the program no longer prints a 'v1' line

Строки accuracy в образце —

  v1     accuracy     0.7700
  v2     accuracy     0.8060   +0.0360

— и шаблон, выбирающий размеченные строки, принял v1 и v2 за метки. То есть живая проверка утверждала, что настоящий прогон печатает строку, начинающуюся с v1. Это верно ровно тогда, когда две версии под тестом оказались пронумерованы 1 и 2.

На ноутбуке они такими и были — запуск одного этого теста обучает ровно две. В CI гоняется весь модуль, TestTrain к тому моменту обучил несколько, и версии были другие.

Фикстура, которая их создаёт, говорит об этом в собственном докстринге:

The version numbers are read back rather than assumed: this module shares one metastore, so how many versions exist by now depends on which other tests have run.

Я его прочитал, процитировал это поведение в тесте, который писал, — и всё равно зашил числа.

figures() теперь выбрасывает строки версий, а их форму проверяет ACCURACY_ROW: сверяет два промежутка пробелов и не интересуется, какие это версии. Прогнано обоими способами — модулем целиком, где версии далеко за двойкой, и этим тестом в одиночку, где они 1 и 2.

The front page opens with a `manage.py diff` run -- agreement, the transitions,
the accuracy pair, fixed against broken, a verdict with a p-value. It is the
first thing anyone reads and the argument for the whole tool, and nothing
checked any part of it.

Three ways it rots, all silent:

A figure gets edited and the others stop following from it. Five hundred rows at
92.0% agreement is forty changed rows; forty changed rows is twenty-two plus
eighteen transitions, and twenty-nine fixed plus eleven broken; 0.7700 to 0.8060
over five hundred rows is eighteen rows, which is twenty-nine minus eleven. All
of that is now asserted, including the verdict's p-value, which is the exact
two-sided McNemar over the discordant pairs and comes to 0.006427 against the
0.006 printed. A sample that contradicts itself would be a poor advertisement
for a tool whose whole pitch is that aggregate numbers hide what moved.

The two READMEs drift apart. It is a transcript of a command, identical in both
languages, so translating it would be inventing it -- they are compared
character for character.

The program stops saying it. A label renamed or a column rewidened leaves the
front page showing output that no longer exists. The existing diff tests assert
that "agreement" and "fixed" appear somewhere, which survives both changes, so
the new test runs the real command over two real trained versions and compares
the shape: every label in the sample must be printed, with its value starting at
the same column, and the labelled-data block must still say "row(s) wrong in"
and "row(s) right in" rather than the "case(s) failing in" the eval path uses.

Nothing was wrong. Every figure already added up and every line already matched,
which is worth saying: this pins a page that was right, rather than fixing one
that was not.

Split where the fixtures are. The two static tests need nothing but the files
and run in half a second; the live one sits beside TestDiff, which already has
the machinery to train a model.

Checked by breaking a copy five ways -- a figure edited, the Russian README left
behind, the p-value adjusted, a label renamed in diff.py, a column widened in
diff.py. Each is caught, and each says which: "agreement and changed disagree",
"verdict says p=0.001, McNemar gives p=0.006427", "the program no longer prints
a 'agreement' line", "'changed' puts its value at column 21, the README shows
column 17".
CI went red on every test job with

    AssertionError: the program no longer prints a 'v1' line

The sample's accuracy rows are

      v1     accuracy     0.7700
      v2     accuracy     0.8060   +0.0360

and the pattern that picks out labelled lines picked up `v1` and `v2` as
labels. So the live check asserted that a real run prints a line beginning
"  v1", which is true only when the two versions under test happen to be
numbered 1 and 2.

On the laptop this was written on they were, because running that one test alone
trains exactly two. In CI the whole module runs, TestTrain has already trained
several by then, and they were not.

The fixture that produces them says this in its own docstring -- "the version
numbers are read back rather than assumed: this module shares one metastore, so
how many versions exist by now depends on which other tests have run" -- and I
read it, quoted its behaviour in the test I was writing, and then hard-coded the
numbers anyway.

`figures()` now drops the version rows, and their shape is checked by
ACCURACY_ROW, which compares the two runs of spacing and does not care which
versions it is looking at. Run both ways since: the module whole, where the
versions are well past two, and that test alone, where they are one and two.
@DenisDrobyshev
DenisDrobyshev merged commit b99bc1a into master Sep 20, 2026
17 checks passed
@DenisDrobyshev
DenisDrobyshev deleted the check/readme-session branch September 20, 2026 19:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant