Skip to content

fix(translate): count the 1500-character cap in UTF-16 code units - #243

Merged
missuo merged 1 commit into
OwO-Network:mainfrom
JounQin:fix/text-limit-utf16
Oct 10, 2026
Merged

missuo merged 1 commit into
OwO-Network:mainfrom
JounQin:fix/text-limit-utf16

Conversation

@JounQin

@JounQin JounQin commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

maxFreeTextLength was checked with utf8.RuneCountInString, but the anonymous oneshot cap is 1500 UTF-16 code units — what String.prototype.length reports in the extension and String.length in the Kotlin iOS client this endpoint is modelled on. A rune count scores an astral character as one where oneshot scores it two, so the guard accepted an oversized request, sent it upstream, and the caller got the generic 503 request failed with status code: 400 instead of the 413 the guard exists to produce.

The check is pre-existing on main (it came in with #220), which is why this is a PR of its own. It was opened against 54158cd; #241 has since merged, so it is rebased onto c9f7467 and now applies the same unit to totalTextLength — the batch total — as well as correcting the unit the README documented.

Evidence

Live single-text probes of the anonymous oneshot endpoint through a local DLX build (auto source, target_lang: EN):

text code points UTF-16 units UTF-8 bytes display columns result
'你'.repeat(800) 800 800 2400 1600 200
'😀'.repeat(750) 750 1500 3000 1500 200
'😀'.repeat(751) 751 1502 3004 1502 400
'a'.repeat(1500) 1500 1500 1500 1500 200
('e' + '\u0301').repeat(1000) 2000 2000 3000 1000 400

Only the UTF-16 column crosses 1500 exactly where the endpoint flips. The accepted 2400-byte, 1600-column CJK request rules out a byte cap and a width cap; the rejected 751-emoji request rules out code points (751) and grapheme clusters (751); the rejected combining request rules out grapheme clusters again (1000 of them). The last row is 2000 runes and so is already over the local cap, so it was probed with the guard’s limit raised to let oneshot answer.

One batched probe shows the cap is the sum and not per text: ["😀".repeat(400), "😀".repeat(400)] is 800 runes and 1600 UTF-16 units — 800 per element — and oneshot answers 400.

Invalid UTF-8 is not a second defect, and needs no validation of its own. encoding/json replaces every invalid byte with U+FFFD before the guard runs (one unit each, exactly what RuneCountInString scored), and json.Marshal would replace them again on the way upstream. A body carrying the truncated sequence f0 90 80 is translated as three replacement characters rather than rejected.

Change

  • textLength counts UTF-16 code units without allocating, and totalTextLength sums it, so the guard measures the whole batch in the unit oneshot counts. unicode/utf8 goes with the old count. BMP text is unaffected — runes and units agree everywhere except a surrogate pair, which is why the rune count looked right.
  • The README’s limit now says UTF-16 code units rather than “Unicode characters (runes)”: a CJK character costs 1, an astral one costs 2.
  • Tests cover both shapes at 1500 and 1501 exactly, astral characters, combining marks, invalid UTF-8 (raw in the body and as surrogate escapes), an empty element inside a batch, and the 413 path, which must return before anything reaches DeepL.

go build ./..., go vet ./... and go test ./... pass. Measured live through a local build: '你'×800 and '😀'×750 answer 200 and 'a'×1501 answers 413, while '😀'×751 changes from oneshot’s 400, surfaced as 503, to 413 text exceeds maximum length: 1502 characters (anonymous oneshot limit is 1500). On the rebased branch the batched ["😀"×400, "😀"×400] likewise answers 413 … 1600 characters locally instead of the same upstream 400.

Fixes #242. Refs #241, #240.

Copilot AI balanced review requested due to automatic review settings October 10, 2026 13:30

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

The anonymous oneshot cap is 1500 UTF-16 code units — what
String.prototype.length reports in the extension and String.length in the
Kotlin iOS client this endpoint is modelled on. utf8.RuneCountInString
scores an astral character as one where oneshot scores it two, so 751
emoji (751 runes, 1502 units) passed the guard, travelled upstream and
came back as the generic 503 "request failed with status code: 400"
instead of the 413 the guard exists to produce.

Count the unit oneshot counts. BMP text is unaffected: runes and UTF-16
units agree everywhere except a surrogate pair, which is why the rune
count looked right.

Invalid UTF-8 is not part of this and needs no validation of its own:
encoding/json has already replaced every invalid byte with U+FFFD before
the guard runs, one unit each, exactly as a rune count scored them.

Measured against the live endpoint through a local build:
- 750 astral characters (1500 units) -> 200, 751 (1502 units) -> 400
- 800 CJK characters (800 units, 2400 bytes, 1600 columns) -> 200
- 1000 combining-mark pairs (2000 units, 1000 graphemes) -> 400
- 1500 ASCII -> 200, 1501 ASCII -> 413 from the guard

Refs OwO-Network#242
@JounQin
JounQin force-pushed the fix/text-limit-utf16 branch from b3001a4 to 10f66b9 Compare October 10, 2026 13:42
@missuo
missuo merged commit 83f97d3 into OwO-Network:main Oct 10, 2026
3 checks passed
@JounQin
JounQin deleted the fix/text-limit-utf16 branch October 10, 2026 17:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(translate): count the 1500-character cap in UTF-16 code units, not runes

3 participants