Держать терминальную сессию в README транскриптом, а не картинкой - #14
Merged
Merged
Conversation
The front page opens with a `manage.py diff` run -- agreement, the transitions, the accuracy pair, fixed against broken, a verdict with a p-value. It is the first thing anyone reads and the argument for the whole tool, and nothing checked any part of it. Three ways it rots, all silent: A figure gets edited and the others stop following from it. Five hundred rows at 92.0% agreement is forty changed rows; forty changed rows is twenty-two plus eighteen transitions, and twenty-nine fixed plus eleven broken; 0.7700 to 0.8060 over five hundred rows is eighteen rows, which is twenty-nine minus eleven. All of that is now asserted, including the verdict's p-value, which is the exact two-sided McNemar over the discordant pairs and comes to 0.006427 against the 0.006 printed. A sample that contradicts itself would be a poor advertisement for a tool whose whole pitch is that aggregate numbers hide what moved. The two READMEs drift apart. It is a transcript of a command, identical in both languages, so translating it would be inventing it -- they are compared character for character. The program stops saying it. A label renamed or a column rewidened leaves the front page showing output that no longer exists. The existing diff tests assert that "agreement" and "fixed" appear somewhere, which survives both changes, so the new test runs the real command over two real trained versions and compares the shape: every label in the sample must be printed, with its value starting at the same column, and the labelled-data block must still say "row(s) wrong in" and "row(s) right in" rather than the "case(s) failing in" the eval path uses. Nothing was wrong. Every figure already added up and every line already matched, which is worth saying: this pins a page that was right, rather than fixing one that was not. Split where the fixtures are. The two static tests need nothing but the files and run in half a second; the live one sits beside TestDiff, which already has the machinery to train a model. Checked by breaking a copy five ways -- a figure edited, the Russian README left behind, the p-value adjusted, a label renamed in diff.py, a column widened in diff.py. Each is caught, and each says which: "agreement and changed disagree", "verdict says p=0.001, McNemar gives p=0.006427", "the program no longer prints a 'agreement' line", "'changed' puts its value at column 21, the README shows column 17".
CI went red on every test job with
AssertionError: the program no longer prints a 'v1' line
The sample's accuracy rows are
v1 accuracy 0.7700
v2 accuracy 0.8060 +0.0360
and the pattern that picks out labelled lines picked up `v1` and `v2` as
labels. So the live check asserted that a real run prints a line beginning
" v1", which is true only when the two versions under test happen to be
numbered 1 and 2.
On the laptop this was written on they were, because running that one test alone
trains exactly two. In CI the whole module runs, TestTrain has already trained
several by then, and they were not.
The fixture that produces them says this in its own docstring -- "the version
numbers are read back rather than assumed: this module shares one metastore, so
how many versions exist by now depends on which other tests have run" -- and I
read it, quoted its behaviour in the test I was writing, and then hard-coded the
numbers anyway.
`figures()` now drops the version rows, and their shape is checked by
ACCURACY_ROW, which compares the two runs of spacing and does not care which
versions it is looking at. Run both ways since: the module whole, where the
versions are well past two, and that test alone, where they are one and two.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Первая страница открывается прогоном
manage.py diff— agreement, переходы, пара accuracy, fixed против broken, вердикт с p-значением. Это первое, что читают, и это аргумент за весь инструмент. Не проверялось ничего.Три способа протухнуть, все молчаливые.
1. Цифру правят, остальные перестают из неё следовать
500 строк при agreement 92.0% — это 40 изменившихся. 40 изменившихся — это 22 + 18 переходов и 29 fixed + 11 broken. 0.7700 → 0.8060 на 500 строках — это 18 строк, то есть 29 − 11.
Теперь это всё утверждается, включая p-значение вердикта: точный двусторонний МакНемар по дискордантным парам даёт 0.006427 против напечатанных
p=0.006.Образец, противоречащий сам себе, — плохая реклама инструменту, чья идея в том, что агрегаты прячут то, что сдвинулось.
2. Два README расходятся
Это транскрипт команды, одинаковый на обоих языках. Перевести его — значит выдумать. Сравниваются символ в символ.
3. Программа перестаёт это печатать
Переименовали метку, расширили колонку — и первая страница показывает вывод, которого больше нет.
Существующие тесты diff утверждают, что
agreementиfixedгде-то встречаются. Это переживает обе поломки. Поэтому новый тест гоняет настоящую команду на двух настоящих обученных версиях и сверяет форму: каждая метка из образца должна печататься, её значение — начинаться в той же колонке, а блок по разметке — по-прежнему говоритьrow(s) wrong in, а неcase(s) failing in, которое использует ветка evals.Ничего не было сломано
Все цифры уже сходились, все строки уже совпадали. Это стоит сказать прямо: PR прибивает страницу, которая была права, а не чинит неправую.
Проверено
Сломал копию пятью способами:
agreement and changed disagreeverdict says p=0.001, McNemar gives p=0.006427diff.pythe program no longer prints a 'agreement' linediff.py'changed' puts its value at column 21, the README shows column 17Разделено по фикстурам: два статических теста не требуют ничего, кроме файлов, и идут полсекунды; живой стоит рядом с
TestDiff, где уже есть машинерия обучения модели.ruff check mlango testsиruff format --check— чисто. 36 тестов в затронутой области проходят.Первый прогон был красный, и это моя ошибка
Упало на всех тестовых job'ах:
Строки accuracy в образце —
— и шаблон, выбирающий размеченные строки, принял
v1иv2за метки. То есть живая проверка утверждала, что настоящий прогон печатает строку, начинающуюся сv1. Это верно ровно тогда, когда две версии под тестом оказались пронумерованы 1 и 2.На ноутбуке они такими и были — запуск одного этого теста обучает ровно две. В CI гоняется весь модуль,
TestTrainк тому моменту обучил несколько, и версии были другие.Фикстура, которая их создаёт, говорит об этом в собственном докстринге:
Я его прочитал, процитировал это поведение в тесте, который писал, — и всё равно зашил числа.
figures()теперь выбрасывает строки версий, а их форму проверяетACCURACY_ROW: сверяет два промежутка пробелов и не интересуется, какие это версии. Прогнано обоими способами — модулем целиком, где версии далеко за двойкой, и этим тестом в одиночку, где они 1 и 2.