Skip to content

2.5.5 regression: strong emphasis adjacent to word/CJK chars fails when content starts or ends with punctuation #688

Description

@AndersonBY

Summary

After upgrading from markdown2 2.5.4 to 2.5.5, strong emphasis sometimes stops parsing when all of these are true:

  • the **...** span is adjacent to alphanumeric or CJK text without surrounding spaces
  • the emphasized content starts or ends with punctuation, for example Chinese quotes or parentheses

On 2.5.4 these cases render correctly. On 2.5.5 the parser leaves literal ** in the output, and in longer inputs it can also emit malformed mixed HTML.

This reproduces with the default parser configuration, no extras required.

Related to #679 because it also looks like an emphasis-regression in the recent parser changes, but this one reproduces without middle-word-em and with plain markdown2.markdown(...).

Minimal repro

import markdown2

cases = [
    'a**“b”**c',
    '**“b”**c',
    'a**(b)**c',
]

print('version:', markdown2.__version__)
for text in cases:
    print('INPUT :', text)
    print('OUTPUT:', markdown2.markdown(text).strip())
    print()

2.5.4

<p>a<strong>“b”</strong>c</p>
<p><strong>“b”</strong>c</p>
<p>a<strong>(b)</strong>c</p>

2.5.5

<p>a**“b”**c</p>
<p>**“b”**c</p>
<p>a**(b)**c</p>

Longer repro that produces malformed mixed HTML

import markdown2

text = '*   **示例**:系统会使用**(方案A)**或**(方案B)**进行处理。'
print(markdown2.markdown(text))

2.5.4

<ul>
<li><strong>示例</strong>:系统会使用<strong>(方案A)</strong><strong>(方案B)</strong>进行处理。</li>
</ul>

2.5.5

<ul>
<li><strong>示例</strong>:系统会使用**(方案A)<strong></strong>(方案B)**进行处理。</li>
</ul>

Suspected regression window

Looking at the source, Markdown._do_italics_and_bold() changed between these versions:

2.5.4

text = self._strong_re.sub(r"<strong>\\2</strong>", text)
text = self._em_re.sub(r"<em>\\2</em>", text)

2.5.5

if not self._iab_processor:
    self._iab_processor = GFMItalicAndBoldProcessor(self, None)
if self._iab_processor.test(text):
    text = self._iab_processor.run(text)

So this looks related to the switch to GFMItalicAndBoldProcessor as the default implementation for _do_italics_and_bold().

Environment

  • Python: 3.14.0
  • markdown2: 2.5.4 vs 2.5.5
  • OS: Windows 11

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions