Repository navigation
fix(translate): count the 1500-character cap in UTF-16 code units - #243
Merged
Merged
Conversation
The anonymous oneshot cap is 1500 UTF-16 code units — what String.prototype.length reports in the extension and String.length in the Kotlin iOS client this endpoint is modelled on. utf8.RuneCountInString scores an astral character as one where oneshot scores it two, so 751 emoji (751 runes, 1502 units) passed the guard, travelled upstream and came back as the generic 503 "request failed with status code: 400" instead of the 413 the guard exists to produce. Count the unit oneshot counts. BMP text is unaffected: runes and UTF-16 units agree everywhere except a surrogate pair, which is why the rune count looked right. Invalid UTF-8 is not part of this and needs no validation of its own: encoding/json has already replaced every invalid byte with U+FFFD before the guard runs, one unit each, exactly as a rune count scored them. Measured against the live endpoint through a local build: - 750 astral characters (1500 units) -> 200, 751 (1502 units) -> 400 - 800 CJK characters (800 units, 2400 bytes, 1600 columns) -> 200 - 1000 combining-mark pairs (2000 units, 1000 graphemes) -> 400 - 1500 ASCII -> 200, 1501 ASCII -> 413 from the guard Refs OwO-Network#242
JounQin
force-pushed
the
fix/text-limit-utf16
branch
from
October 10, 2026 13:42
b3001a4 to
10f66b9
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
maxFreeTextLengthwas checked withutf8.RuneCountInString, but the anonymous oneshot cap is 1500 UTF-16 code units — whatString.prototype.lengthreports in the extension andString.lengthin the Kotlin iOS client this endpoint is modelled on. A rune count scores an astral character as one where oneshot scores it two, so the guard accepted an oversized request, sent it upstream, and the caller got the generic503 request failed with status code: 400instead of the413the guard exists to produce.The check is pre-existing on
main(it came in with #220), which is why this is a PR of its own. It was opened against54158cd; #241 has since merged, so it is rebased ontoc9f7467and now applies the same unit tototalTextLength— the batch total — as well as correcting the unit the README documented.Evidence
Live single-text probes of the anonymous oneshot endpoint through a local DLX build (auto source,
target_lang: EN):text'你'.repeat(800)'😀'.repeat(750)'😀'.repeat(751)'a'.repeat(1500)('e' + '\u0301').repeat(1000)Only the UTF-16 column crosses 1500 exactly where the endpoint flips. The accepted 2400-byte, 1600-column CJK request rules out a byte cap and a width cap; the rejected 751-emoji request rules out code points (751) and grapheme clusters (751); the rejected combining request rules out grapheme clusters again (1000 of them). The last row is 2000 runes and so is already over the local cap, so it was probed with the guard’s limit raised to let oneshot answer.
One batched probe shows the cap is the sum and not per text:
["😀".repeat(400), "😀".repeat(400)]is 800 runes and 1600 UTF-16 units — 800 per element — and oneshot answers 400.Invalid UTF-8 is not a second defect, and needs no validation of its own.
encoding/jsonreplaces every invalid byte with U+FFFD before the guard runs (one unit each, exactly whatRuneCountInStringscored), andjson.Marshalwould replace them again on the way upstream. A body carrying the truncated sequencef0 90 80is translated as three replacement characters rather than rejected.Change
textLengthcounts UTF-16 code units without allocating, andtotalTextLengthsums it, so the guard measures the whole batch in the unit oneshot counts.unicode/utf8goes with the old count. BMP text is unaffected — runes and units agree everywhere except a surrogate pair, which is why the rune count looked right.go build ./...,go vet ./...andgo test ./...pass. Measured live through a local build:'你'×800and'😀'×750answer 200 and'a'×1501answers 413, while'😀'×751changes from oneshot’s 400, surfaced as503, to413 text exceeds maximum length: 1502 characters (anonymous oneshot limit is 1500). On the rebased branch the batched["😀"×400, "😀"×400]likewise answers413 … 1600 characterslocally instead of the same upstream 400.Fixes #242. Refs #241, #240.