fix(content)!: spell literal text so python-markdown reads it as text (0.6.0) - #39
Merged
Merged
Conversation
from_transport handed markdownify's output on as Markdown, and markdownify escapes nothing it was not asked to: text a user typed into GLPI came back as syntax once any peer rendered it. Measured through this reader into python-markdown with the four extensions to_transport uses: \\serveur lost a backslash, __init__ and ______ became emphasis, a "-----------" line under text made a heading, "* point" and "> merci" lines made a list and a quote, "#4521" at a line start made a heading, "[1]: url" was consumed as a reference definition, and a | in a cell dropped the rest of the row. The converter is now a MarkdownConverter subclass. escape() escapes nothing itself: it stands each context-dependent character in for with a private-use code point, and each container python-markdown parses on its own -- paragraph, div, list item, quote, heading, cell, document -- decides every stand-in in it at once, with its whole Markdown in view. The rules replay python-markdown's own: its code-span pattern, its escape set (asked of the renderer, not copied), its bracket counting, its emphasis pairing including the NOT_STRONG exemption, its rule, setext, fence, list, quote and reference-definition line tests, and its e-mail autolink, which it matches only after code spans, escapes, links and images are stashed. A character is escaped only where python-markdown would misread it, with a backslash it removes again, or with a character reference where no backslash works (< & = ~): ordinary prose -- fichier_de_test_v2.xlsx, C:\Temp, R&D, a # or a - mid-sentence -- comes back byte for byte as before. Structure markdownify garbled on the way through is fixed where literal safety depended on it: continuation lines indented by four spaces, so nested lists nest and keep their numbers; content after a nested list starts a block of its own instead of joining the list's last item; a quote in an item gets the blank line python-markdown needs, and a quote opening an item keeps its later lines lazy; items python-markdown writes loose are written loose, so the second read equals the first; a list on a line with three bullets is spaced, because python-markdown cannot nest anything under such a line; a <pre> in an item or quote is an indented code block, except where python-markdown reads none, where its lines are kept as literal text. Script, style and title bodies are dropped instead of leaking into the text, <s>/<del>/<strike> keep their words without a ~~ python-markdown would display, images stay images in headings and cells, and a source newline is a space, as HTML displays it. The HTML standard's obsolete elements now make a body HTML. The degraded path spells its stripped text through the same rules, so a body too deep to convert is literal-safe too. test_literal_text.py holds the property -- to_transport(from_transport( html)) displays what html displays, compared with an HTML parser, and the Markdown is a fixed point -- over the measured misreadings, one case per construct, realistic bodies and a seeded fuzzer (8 seeds x 50 bodies in the suite; 30,000 bodies across 600 seeds measured clean). What the format cannot carry is an inventory of strict xfails, each asserted to stabilise after one cycle. The nested-list round-trip loss is gone from the corpus. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Four passes of the literal-text reader cost quadratic time on a body dense with one character, measured on 20,000 repetitions: unclosed brackets (settle_brackets counted forward from each [), words opening with an underscore (the pairing check compared every opener with every later run), backtick runs (_code_spans rebuilt the whole text after every span) and a long numbered list (each bullet counted its previous siblings, which markdownify's own converter did too -- 34 s before this change). They took 8 to 117 seconds each and take well under one now: - brackets are matched right to left against a stack of unclosed ], an escaped [ handing its ] back, which is python-markdown's own count; - underscore pairing is one pass remembering whether an opener was seen; - _code_spans searches on from each span's end, rebuilding the text only for the escaped-backslash run python-markdown consumes on its own; - bullets are numbered once per list, and a list's shown items are counted once. A complexity test holds each of the four under a ten-second budget. Mutation testing, in a disposable worktree, over 65 mutants of the reader's guards left eight alive before this commit. Six were test gaps, now closed: a backtick touching a generated code span, an escaped bracket in link text that python-markdown does not count, seven underscores mid-sentence, asterisks in an item and its nested item, a style block between a nested list and a code block, and a line break in a cell or a heading. Two were dead code, now gone: the nested list's blank line no longer checks for a following sibling -- the item strips it when nothing follows -- and from_transport no longer strips a result the document converter has already stripped, before settling it, which is where the strip matters. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Literal-safe Markdown on the read path (588e97d, 7ea833a), and the markdownify floor it needs: >=1.2, since the converter subclasses MarkdownConverter and relies on the 1.x hooks -- escape(text, parent_tags), convert_*(el, text, parent_tags), convert__document_ -- which 0.13 does not have. 1.2.2 and 1.2.3 are the versions measured. The changelog records the fixes, the two behaviour changes a caller can see (a value holding one real HTML element is read as HTML throughout, write models included; Markdown read from GLPI changes once for any body the old reader let through as syntax) and the inventory of what Markdown cannot carry. The user guide and the API reference say what the escaping is and is not, with an example checked against the code, and the two skills that describe .content now say not to strip the backslashes. A correction to 7ea833a's message: nine of the 65 mutants survived before it, not eight, and seven were test gaps -- the seventh is alt text over several lines, which a blank line in it split out of its paragraph once the collapse was removed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The deepest case was 300 levels. CPython 3.10, which CI runs, spends about three frames per level where 3.12 spends two, so its cliff is about 328 levels from a shallow stack -- measured by easyvista_python_client's port of this converter, 2026-09-30 -- and 300 left 14 levels of margin under pytest. 250 still proves the point, being past the old fixed bound of 200. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
8 tasks
baraline
marked this pull request as ready for review
October 1, 2026 07:49
|
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #39 +/- ##
==========================================
+ Coverage 97.59% 97.64% +0.04%
==========================================
Files 90 90
Lines 3163 3313 +150
==========================================
+ Hits 3087 3235 +148
- Misses 76 78 +2 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
from_transportpassed markdownify's output on as Markdown with escaping switched off. Literal text in a GLPI body therefore changed meaning as soon as python-markdown rendered it, which is whatto_transportand any consumer does.Measured through the reader into python-markdown, with the four extensions
to_transportuses:\\serveur\comptalost a backslash, and__init__became bold;-----line under text made a heading;* pointand> mercilines became a list and a quote;#4521at the start of a line became a heading;[1]: https://…was consumed as a reference definition;|inside a table cell dropped the rest of the row;What changes
GlpiContentConverter.from_transport, and so every read model's.content) escapes a character only where python-markdown would read it as syntax:<,&,=,~) where no backslash escape exists.fichier_de_test_v2.xlsx,C:\Temp,R&Dor a#in mid-sentence;<https://…>;<script>,<style>and<title>bodies are dropped, where they used to leak as prose;font,center,strike, …) are recognised;<s>and<del>keep their words without a literal~~.markdownify>=1.2, because the subclass needs the 1.x hooks. Version 0.6.0.Breaking for consumers
.contentnow carries backslash escapes and character references wherever literal text would otherwise be misread. Render it before displaying or indexing it..contenttherefore change once.Verification
pytest -m "not integration"with coverage: 1569 passed, 17 xfailed, 97.6 % (the gate is 95 %).ruff,mypy(strict),unasync_build.py --checkandsphinx -Wpass. Sphinx was built offline, with intersphinx disabled locally. The wheel was checked for test files.to_transport(from_transport(html))displays whathtmldisplays. 50,000 fuzzed bodies gave 0 failures.Still to do before this leaves draft
These come from that review.
defeats the line rules in three places: the start of the first block, the start of a list item, and the end of the last block.SRV___PROD___01;a _ b _ c.The EasyVista port of this converter: baraline/easyvista_python_client#6
🤖 Generated with Claude Code