Skip to content

fix(filters): separate each match on its own line for source: watches - #4492

Open
Yi-111-a wants to merge 1 commit into
dgtlmoon:masterfrom
Yi-111-a:fix/source-type-filter-match-separator
Open

Yi-111-a wants to merge 1 commit into
dgtlmoon:masterfrom
Yi-111-a:fix/source-type-filter-match-separator

Conversation

@Yi-111-a

Copy link
Copy Markdown

Fixes #4477

The problem

A watch on a source: URL passes append_pretty_line_formatting=False to the filter helpers, and its filtered text is then kept verbatim instead of being run through html_to_text():

# changedetectionio/processors/text_json_diff/processor.py
append_pretty_line_formatting=not self.watch.is_source_type_url,
...
if watch.is_source_type_url:
    # For source URLs, keep raw content
    stripped_text = html_content

TEXT_FILTER_LIST_LINE_SUFFIX is "<br>", which only works as a line separator because Inscriptis turns it back into a newline later on. For a source: watch that step never runs, so the code skipped the separator entirely — and with no separator at all, every match ran straight into the next on one line.

Reported on source:https://www.wikipedia.org/ with the //link/@href filter:

Current:
/static/apple-touch/wikipedia.png/static/favicon/wikipedia.ico//creativecommons.org/licenses/by-sa/4.0///upload.wikimedia.orghttps://wikis.world/@wikipedia

The fix

filter_match_separator() picks a real "\n" instead of "<br>" when the result will not be passed through Inscriptis:

if not append_pretty_line_formatting:
    return "\n" if len(html_block) else ""

xpath_filter(), xpath1_filter() and include_filters() all carried the same if append_pretty_line_formatting and len(html_block) and ... condition, so all three now call the helper. That also fixes the CSS and xpath1: variants of the same bug, which had the identical symptom.

Two details preserved:

  • The first match still never gets a separator, so no leading newline is introduced.
  • The Inscriptis path is unchanged, byte for byte — same <br> marker, same skip for br/hr/div/p (which already break a line).

Result for the reported case:

/static/apple-touch/wikipedia.png
/static/favicon/wikipedia.ico
//creativecommons.org/licenses/by-sa/4.0/
//upload.wikimedia.org
https://wikis.world/@wikipedia

Tests

New changedetectionio/tests/unit/test_filter_match_separator.py covers the multi-match case for XPath, xpath1: and CSS, the single-match and no-match cases (no stray leading newline), and the helper itself for both paths.

changedetectionio/tests/unit/ + test_xpath_selector_unit.py   467 passed
test_css_selector, test_xpath_selector, test_xpath_default_namespace,
test_source, test_lxml_concurrency, test_filter_exist_changes,
test_ignore_text                                             44 passed

dgtlmoon#4477

A `source:` watch passes append_pretty_line_formatting=False to the filter
helpers, and then keeps the filtered text verbatim instead of running it
through html_to_text(). TEXT_FILTER_LIST_LINE_SUFFIX ('<br>') only works
because Inscriptis turns it into a newline, so for that path no separator was
emitted at all and every match ran into the next on a single line.

Add filter_match_separator(), which picks '\n' instead of '<br>' when the
result will not be passed through Inscriptis, and use it from
xpath_filter(), xpath1_filter() and include_filters() - all three carried the
same condition. The first match still never gets a separator, so no leading
newline is introduced, and the Inscriptis path is byte-for-byte unchanged.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

XPath multiple match returns one line instead multiple lines for URLs prefixed with "source:"

1 participant