Repository navigation
[Performance] Render layout masks within polygon bounds - #5217
Merged
Merged
Conversation
changdazhou
approved these changes
Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Layout visualization currently allocates a full-page mask for every polygon, stacks all masks, and scans each full page to locate covered pixels. Documents with many layout regions spend substantial CPU time and memory on this work.
This change rasterizes each polygon into its clipped bounding rectangle and blends that region in place. It preserves the existing class colors, sequential overlap blending, labels, and reading-order annotations. The unused full-page mask helper is removed and its OpenCV dependency guard moves to
draw_mask.The change is limited to
paddlex/inference/models/layout_analysis/result.py. It uses the existingLayoutAnalysisResult._to_imgcontract, where the mask canvas has the same dimensions as the image.Validation
CPU comparison against the unmodified
developimplementation, using OpenCV 4.10.0, NumPy 2.3.5, and Pillow 12.3.0:git diff --check.End-to-end ablation against unmodified develop
Fresh measurements on 2026-10-08 compare develop (
c50f5da) with this PR (7ed9604). The frontend trees differ only in the mask-rendering file. Both use the same original backend and HTTPX JSON client, without the earlier queue, visual batching, RoPE, or JSON optimizations.One A800 80 GB, PaddleOCR-VL-1.6 with PP-DocLayoutV3, vLLM 0.26.0; all image outputs enabled. Each configuration starts fresh services, warms up on the same 16 pages, then measures the same 50 pages three times. Prefix and multimodal processor caches are disabled on both sides.
Values are medians. Timing includes input reading/encoding, HTTP processing and transport, response parsing, and saving full responses. Startup, warmup, and the common one-time TIFF preparation are excluded.
All 12 measured runs completed: 600 pages, zero failed or missing pages. All 2,364 returned image assets match the fresh develop reference in encoded bytes and decoded pixels, with no missing Markdown image references. Layout boxes match for every page. Repeated runs have some Markdown text differences in both baseline and candidate; image equality is not a claim of full OCR quality equivalence.
Source, input, and model hashes were verified before and after the experiment. Each configuration processed 2,967 backend region requests including warmup, with the same prompt-token total and the same termination counts.