Skip to content

[Performance] Render layout masks within polygon bounds - #5217

Merged
scyyh11 merged 1 commit into
PaddlePaddle:developfrom
scyyh11:perf/local-layout-mask
Oct 8, 2026
Merged

scyyh11 merged 1 commit into
PaddlePaddle:developfrom
scyyh11:perf/local-layout-mask

Conversation

@scyyh11

@scyyh11 scyyh11 commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Layout visualization currently allocates a full-page mask for every polygon, stacks all masks, and scans each full page to locate covered pixels. Documents with many layout regions spend substantial CPU time and memory on this work.

This change rasterizes each polygon into its clipped bounding rectangle and blends that region in place. It preserves the existing class colors, sequential overlap blending, labels, and reading-order annotations. The unused full-page mask helper is removed and its OpenCV dependency guard moves to draw_mask.

The change is limited to paddlex/inference/models/layout_analysis/result.py. It uses the existing LayoutAnalysisResult._to_img contract, where the mask canvas has the same dimensions as the image.

Validation

CPU comparison against the unmodified develop implementation, using OpenCV 4.10.0, NumPy 2.3.5, and Pillow 12.3.0:

  • 50 saved real document pages: complete rendered images are pixel-identical.
  • 1,515 synthetic cases: blended pixels and complete images are identical. Cases cover clipping, concave and self-intersecting polygons, degenerate polygons, fractional coordinates, repeated classes, overlaps, small images, and smaller mask canvases.
  • 15,000 additional rasterization comparisons, including coordinates from -4096 to 4095: identical blended pixels.
  • A 21 × 21 polygon region on an 800 × 600 image allocates a 21 × 21 mask.
  • Black, Flake8, isort, license headers, dependency-import checks, Python 3.8 annotation checks, and git diff --check.

End-to-end ablation against unmodified develop

Fresh measurements on 2026-10-08 compare develop (c50f5da) with this PR (7ed9604). The frontend trees differ only in the mask-rendering file. Both use the same original backend and HTTPX JSON client, without the earlier queue, visual batching, RoPE, or JSON optimizations.

One A800 80 GB, PaddleOCR-VL-1.6 with PP-DocLayoutV3, vLLM 0.26.0; all image outputs enabled. Each configuration starts fresh services, warms up on the same 16 pages, then measures the same 50 pages three times. Prefix and multimodal processor caches are disabled on both sides.

Requests develop (s) Local masks (s) Time reduction Throughput
50 sequential single-page requests 114.590 88.460 22.8% 1.30x
One 50-page request 72.658 45.978 36.7% 1.58x

Values are medians. Timing includes input reading/encoding, HTTP processing and transport, response parsing, and saving full responses. Startup, warmup, and the common one-time TIFF preparation are excluded.

All 12 measured runs completed: 600 pages, zero failed or missing pages. All 2,364 returned image assets match the fresh develop reference in encoded bytes and decoded pixels, with no missing Markdown image references. Layout boxes match for every page. Repeated runs have some Markdown text differences in both baseline and candidate; image equality is not a claim of full OCR quality equivalence.

Source, input, and model hashes were verified before and after the experiment. Each configuration processed 2,967 backend region requests including warmup, with the same prompt-token total and the same termination counts.

@scyyh11
scyyh11 merged commit 2f78d8f into PaddlePaddle:develop Oct 8, 2026
4 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants