Skip to content

Add DecompressTo for zero-alloc output buffer reuse - #9

Merged
ChillerDragon merged 3 commits into
teeworlds-go:masterfrom
jxsl13:decompress-to
Jun 14, 2026
Merged

ChillerDragon merged 3 commits into
teeworlds-go:masterfrom
jxsl13:decompress-to

Conversation

@jxsl13

@jxsl13 jxsl13 commented Jun 13, 2026

Copy link
Copy Markdown
Contributor

What

Adds Huffman.DecompressTo(dst, data []byte) ([]byte, error) which appends the
decompressed bytes into a caller-supplied buffer and returns the extended slice
(same semantics as the builtin append). Decompress becomes a thin wrapper, so
existing behavior is unchanged.

Why

Decompress allocates a fresh []byte{} and append-grows it on every call
(~2 allocs / ~192 B for a typical snapshot payload). On hot paths that
decompress a continuous packet stream — e.g. a 50 Hz Teeworlds snapshot feed —
that is pure GC churn for output that the caller often copies out or processes
immediately anyway.

By passing buf[:0], callers reuse the backing array across calls and reach a
steady state of 0 allocs/op:

BenchmarkDecompressTo-12   258517   5731 ns/op   0 B/op   0 allocs/op

Notes

  • huff is not mutated, so a single Huffman value remains safe for concurrent
    DecompressTo calls with distinct dst buffers.
  • dst may be nil.
  • Added TestDecompressToMatchesDecompress (byte-identical to Decompress,
    proves a stale reused buffer is reset) and BenchmarkDecompressTo.

🤖 Generated with Claude Code

jxsl13 and others added 2 commits June 13, 2026 18:40
Huffman.Decompress allocates a fresh output slice and append-grows it on
every call (~2 allocs / ~192 B per call for a typical snapshot payload).
On hot paths that decompress a continuous packet stream (e.g. a 50 Hz
snapshot feed) that is pure GC churn.

DecompressTo(dst, data) appends the decompressed bytes into a
caller-supplied buffer and returns the extended slice (append semantics).
Passing buf[:0] reuses the backing array, reaching a steady state of
0 allocs/op. Decompress is now a thin wrapper, so behavior is unchanged.

BenchmarkDecompressTo: 5731 ns/op  0 B/op  0 allocs/op

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
b.Loop() is the modern benchmark idiom: no manual b.N loop or
ResetTimer, and it keeps the loop body's inputs/results live so the
compiler cannot eliminate the decompression work being measured.

This requires the go directive to be >= 1.24, so bump go 1.22.3 -> 1.24.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@jxsl13

jxsl13 commented Jun 13, 2026

Copy link
Copy Markdown
Contributor Author

Updated the benchmark to use for b.Loop() (the modern Go 1.24 idiom — no manual b.N/ResetTimer, and it keeps inputs/results live so the work can't be elided). This requires bumping the module go directive 1.22.3 → 1.24.0, included in the latest commit. Let me know if you'd rather keep the lower minimum Go version and I'll revert to the classic b.N loop.

The install snippet is a shell block but used // comments, which
shellcheck (via the README lintdown.sh CI check) rejects with SC1127.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@jxsl13

jxsl13 commented Jun 13, 2026

Copy link
Copy Markdown
Contributor Author

@ChillerDragon i need a new version v2.1.0 :O

@ChillerDragon

Copy link
Copy Markdown
Member

Ok great this is becoming a slop project 🚀

Supporting older versions always great but sure we can upgrade it.

@ChillerDragon
ChillerDragon merged commit 747cb23 into teeworlds-go:master Jun 14, 2026
1 check passed
jxsl13 added a commit to jxsl13/huffman that referenced this pull request Aug 16, 2026
… perf-improvements

Resolves huffman.go: master added DecompressTo (PR teeworlds-go#9), this branch
rewrote the decoder. Kept the optimized decoder body and moved it under
the new DecompressTo signature.

Two things the textual merge could not have caught:

  - the optimized body allocated its own output slice, which would have
    shadowed the new dst parameter and silently dropped DecompressTo's
    append semantics
  - its output-size estimate has an 8 KiB floor. Requiring the caller's
    buffer to clear that floor reallocated on every call, so the
    benchmark's 4 KiB buffer never reached the 0 allocs/op the API
    exists for. Pre-growing is now gated on the compressed size instead.

DecompressTo: 5731 ns/op -> 1857 ns/op, still 0 allocs/op.

Adds corpus-wide coverage for the append contract, buffer reuse and the
zero-allocation property.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants