zstd

package
v1.3.3 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 29, 2026 License: Apache-2.0 Imports: 7 Imported by: 0

README

compress/zstd

compress/zstd is an RFC 8878 package for TinyGo and Go web servers. It streams a valid Content-Encoding: zstd representation and calculates its SHA-256 digest during output, so cache entries can retain the encoded bytes and a strong ETag without hashing the bytes in a second pass. It also decodes zstd from any encoder; see Decoding.

encoded, result, err := zstd.EncodeAll(body)
if err != nil {
	return err
}
header.Set("Content-Encoding", zstd.ContentEncoding)
header.Set("ETag", result.ETag())

ETag calculation is enabled by default. Disable it for responses such as Cache-Control: no-store; this avoids allocating and updating SHA-256:

encoded, result, err := zstd.EncodeAll(body, zstd.WithETag(false))
// result.ETagEnabled is false, result.ETag() is empty, and SHA256 is zero.

NewWriter provides the bounded streaming form. Call Close, then Result; closing the encoder does not close its destination.

Constructing an encoder writes nothing to the destination. As in compress/gzip, the frame header goes out with the first Write, Flush, or Close. A handler can therefore wrap its http.ResponseWriter before rendering and still answer a rendering failure with an uncompressed error response: nothing has reached the wire, so the status is not committed and Content-Encoding can still be dropped.

Flush emits the buffered input as complete blocks so a reader can decode everything written so far, which is what streaming responses and server-sent events need between chunks. It neither ends the frame nor flushes the destination, so flush the destination separately:

if _, err := z.Write(chunk); err != nil {
	return err
}
if err := z.Flush(); err != nil {
	return err
}
w.(http.Flusher).Flush()

Flushing before a block fills reduces the compression ratio, so flush per chunk rather than per Write.

Reset starts a new frame on a new destination, so a server can pool encoders across responses instead of building one per response. It keeps what the encoder is made of — under TinyGo a 128 KiB block buffer and a 16 KiB match table, which is nearly all of its footprint — and keeps the WithETag setting chosen at NewWriter. As with NewWriter, nothing reaches the new destination until the caller writes, flushes, or closes.

z := pool.Get().(*zstd.Writer)
z.Reset(w)
defer func() { z.Close(); pool.Put(z) }()

fasthttp does exactly this: it is the encoder TinyGo builds of that fork compress with, because klauspost's decoder is assembly TinyGo cannot link.

Decoding

DecodeAll decodes a whole body; Reader decodes a stream a block at a time. Both accept every frame RFC 8878 defines except those that need a dictionary, skip skippable frames, run on across concatenated frames, and verify content checksums when a frame carries one.

body, err := zstd.DecodeAll(nil, encoded, zstd.WithMaxOutput(10<<20))

Input from the network needs two limits, and the decoder applies both:

  • Window. A frame declares how much earlier content its matches may reach, and a decoder has to keep that much. A frame declaring more than WithMaxWindow fails with ErrWindowTooLarge before anything is decoded. The default is 8 MiB, the limit RFC 9659 sets for the zstd content coding and within which the reference CLI stays up to level 19. A Reader holds at most twice the window plus one 128 KiB block, so this bounds its memory.
  • Output. A few kilobytes of zstd can decode to gigabytes. WithMaxOutput fails with ErrOutputTooLarge once the content passes a length; a Reader returns the content up to it first. There is no default, so set it for untrusted input.

NewReader reads nothing until the first Read, so it fails only on an invalid option. Reset keeps a Reader's buffers for pooling, and Close releases them without closing the underlying reader. A pool can build its Readers without a stream and hand them one with Reset:

var readers = sync.Pool{New: func() any {
	r, _ := zstd.NewReader(nil, zstd.WithMaxOutput(10<<20))
	return r
}}

r := readers.Get().(*zstd.Reader)
defer readers.Put(r)
if err := r.Reset(body); err != nil {
	return err
}
_, err := io.Copy(dst, r)

Failures are reported the same way by both implementations:

error meaning
io.ErrUnexpectedEOF the input ends inside a frame
ErrCorrupt anything else malformed, including a checksum mismatch; the wrapping error says where
ErrWindowTooLarge the frame's window exceeds WithMaxWindow
ErrOutputTooLarge the content exceeds WithMaxOutput
ErrDictionaryRequired the frame was compressed against a dictionary

A read error from the underlying stream is returned as it is.

The TinyGo decoder is as strict as the reference implementation where klauspost is lenient: an entropy-coded bitstream must end exactly where its last symbol does, and a block may not exceed the smaller of the window and 128 KiB. The two implementations can therefore disagree on a damaged frame that no encoder would write, but not on anything an encoder does write.

Measured on an Apple M-series machine, over the reference CLI's frames of this package's own sources:

decoder throughput
TinyGo decoder, host Go 270–450 MB/s
TinyGo decoder, TinyGo 0.42 ~250 MB/s
klauspost through this API, host Go 510–960 MB/s

Neither allocates per call once its pools are warm. Under TinyGo the decoder adds about 92 KB to a program that uses DecodeAll and Reader, and nothing to one that only encodes.

Implementation selection

  • normal host Go builds use github.com/klauspost/compress/zstd
  • TinyGo builds use this package's bounded pure-Go encoder and decoder
  • go build -tags force_tinygo_logic forces the TinyGo-compatible code on host Go

Both implementations expose the same Writer, Result, Option, EncodeAll, Reader, DecoderOption and DecodeAll API, Reset included. Encoded bytes and therefore ETags may differ between implementations.

The host backend uses the klauspost default compression level with one encoder, a 128 KiB window, lower-memory mode, and no frame checksum. The TinyGo backend has the following supported subset:

  • standard Zstandard frames with a 128 KiB window
  • raw and RLE blocks of at most 128 KiB, including profitable interior runs
  • compressed blocks carrying many sequences, from a greedy matcher that keeps one candidate per hash slot
  • FSE sequence tables fitted to each block, falling back to the format's predefined tables when a block has too few sequences to pay for a description, and to RLE tables when a stream carries one symbol
  • repeat offsets, for the common case of a match at the previous distance
  • a lazy step, which defers a match by one byte when the next position starts a longer one
  • Huffman-coded literals, in one stream or four, with the direct weight representation; raw and RLE literal blocks are used where either is smaller
  • streaming output with at most one input block retained
  • Flush at block boundaries without ending the frame
  • SHA-256 and encoded size calculated over bytes successfully written
  • strong, quoted ETag formatting for the encoded representation

Every block falls back to raw or RLE when a compressed one would not be smaller, so output never exceeds the input by more than the block headers.

Compression ratio

Measured against compress/flate at its default level, which is the encoding a server would otherwise negotiate:

payload this encoder deflate
14 KiB HTML listing 8.2% 11.6%
11 KiB JSON array 11.0% 13.3%
5 KiB varied text 29.6% 26.6%
one repeated string 1.4% 1.6%
incompressible 100.1% 100.1%

Varied prose is the one case that loses, and the breakdown says why: its cost is 1247 bytes of sequences against 233 of literals, where deflate is finding word-level repeats this matcher does not.

TestRatioAgainstDeflate holds these within a stated multiple of deflate, and every case in the suite decodes through the reference implementation, so no ratio here was bought with bytes a real decoder would reject.

Matching stays inside the current block, which is what bounds memory to one retained block. A match therefore never reaches back into an earlier block, even though the window would allow it, so a payload whose repeats are further apart than 128 KiB compresses worse than a general-purpose encoder would manage.

Public API exclusions

  • dictionaries and the seekable format, for encoding and decoding
  • compression-level or dictionary options
  • writing frame content checksums (the cache digest is separate); the decoder verifies them

The TinyGo backend additionally omits unsafe code, assembly, and CGo, and writes Huffman weights only in the direct representation, never FSE-compressed. That representation encodes its weight count as 127 plus it, so the largest literal byte in a block must be 128 or below; a block whose literals reach higher stores them instead of coding them. Binary payloads therefore compress through their matches alone.

Documentation

Overview

Package zstd writes RFC 8878 Zstandard frames, calculating cache metadata over the encoded representation while it is emitted, and reads them back.

Host Go uses github.com/klauspost/compress/zstd. TinyGo uses this package's own bounded encoder and decoder, which can also be selected on host Go with the shared force_tinygo_logic build tag. Both implementations expose the same API. The encoders calculate Result's SHA-256 digest over bytes successfully written and write nothing to the destination until the caller writes, flushes, or closes. The decoders accept any frame without a dictionary, report failures with the same errors, and refuse windows beyond WithMaxWindow, 8 MiB by default.

Writer and Reader are poolable through Reset, which is how the fasthttp fork in this repository compresses responses under TinyGo.

Index

Examples

Constants

View Source
const ContentEncoding = "zstd"

ContentEncoding is the HTTP content-coding token for Zstandard.

Variables

View Source
var (
	// ErrClosed reports use of a Writer or Reader after Close.
	ErrClosed            = errors.New("zstd: use after close")
	ErrResultUnavailable = errors.New("zstd: result is unavailable before a successful close")
)
View Source
var (
	// ErrCorrupt reports input that is not valid Zstandard: a bad magic
	// number, a malformed block, a match reaching outside the output, a
	// content size or checksum that does not match. Errors that wrap it say
	// where the input went wrong. Input that simply ends early is reported as
	// io.ErrUnexpectedEOF instead.
	ErrCorrupt = errors.New("zstd: corrupt input")

	// ErrWindowTooLarge reports a frame that declares a window larger than
	// the decoder accepts; see WithMaxWindow. It is checked before the frame
	// is decoded, so it costs no memory.
	ErrWindowTooLarge = errors.New("zstd: frame window exceeds the limit")

	// ErrOutputTooLarge reports content longer than WithMaxOutput allows.
	ErrOutputTooLarge = errors.New("zstd: decoded content exceeds the limit")

	// ErrDictionaryRequired reports a frame compressed against a dictionary,
	// which this package does not support.
	ErrDictionaryRequired = errors.New("zstd: frame requires a dictionary")
)

Functions

func DecodeAll added in v1.3.3

func DecodeAll(dst, src []byte, options ...DecoderOption) ([]byte, error)

DecodeAll decodes every frame in src and appends the content to dst. Skippable frames are skipped, and empty input decodes to nothing. On error, what the returned slice holds beyond dst is unspecified.

Example
package main

import (
	"fmt"

	"github.com/shibukawa/tinygodriver/compress/zstd"
)

func main() {
	encoded, _, err := zstd.EncodeAll([]byte("request body"))
	if err != nil {
		panic(err)
	}
	// Bound what a few bytes of untrusted input may expand to.
	body, err := zstd.DecodeAll(nil, encoded, zstd.WithMaxOutput(1<<20))
	if err != nil {
		panic(err)
	}
	fmt.Println(string(body))
}
Output:
request body

Types

type DecoderOption added in v1.3.3

type DecoderOption interface {
	// contains filtered or unexported methods
}

DecoderOption configures DecodeAll and NewReader.

func WithMaxOutput added in v1.3.3

func WithMaxOutput(n int64) DecoderOption

WithMaxOutput limits the decoded content to n bytes across all frames; longer content fails with ErrOutputTooLarge, after a Reader has returned the first n bytes. Zero, the default, sets no limit. Without one, a few kilobytes of input can decode to gigabytes, so set it for input from untrusted sources.

func WithMaxWindow added in v1.3.3

func WithMaxWindow(n int) DecoderOption

WithMaxWindow sets the largest window a frame may declare, from 1 KiB to 1 GiB. The default is 8 MiB, the limit RFC 9659 sets for HTTP. A Reader keeps up to twice the window in memory, so this is what bounds its footprint; a larger frame fails with ErrWindowTooLarge before it is decoded.

type Option

type Option interface {
	// contains filtered or unexported methods
}

Option configures both the host-Go and TinyGo encoders.

func WithETag

func WithETag(enabled bool) Option

WithETag controls whether SHA-256 cache metadata is calculated while the encoded representation is written. It is enabled by default. When disabled, Result.SHA256 is zero and Result.ETag returns an empty string.

type Reader added in v1.3.3

type Reader struct {
	// contains filtered or unexported fields
}

Reader decompresses a stream of Zstandard frames, skipping skippable ones. Reader is not safe for concurrent use.

func NewReader added in v1.3.3

func NewReader(r io.Reader, options ...DecoderOption) (*Reader, error)

NewReader returns a Reader that decompresses r. Unlike compress/gzip's NewReader it reads nothing until the first Read, so it cannot fail on the stream's content; it fails only on an invalid option. r may be nil, for a Reader that a pool will Reset before use; reading it first fails.

Example
package main

import (
	"bytes"
	"fmt"
	"io"

	"github.com/shibukawa/tinygodriver/compress/zstd"
)

func main() {
	encoded, _, err := zstd.EncodeAll([]byte("streamed body"))
	if err != nil {
		panic(err)
	}
	r, err := zstd.NewReader(bytes.NewReader(encoded))
	if err != nil {
		panic(err)
	}
	defer r.Close()
	body, err := io.ReadAll(r)
	if err != nil {
		panic(err)
	}
	fmt.Println(string(body))
}
Output:
streamed body

func (*Reader) Close added in v1.3.3

func (z *Reader) Close() error

Close releases the Reader's decoder. It does not close the underlying reader. Read reports ErrClosed until Reset.

func (*Reader) Read added in v1.3.3

func (z *Reader) Read(p []byte) (int, error)

Read decompresses into p. It returns io.EOF at the end of the last frame, io.ErrUnexpectedEOF if the stream ends inside one, and ErrOutputTooLarge once it has returned as much content as WithMaxOutput allows and more remains.

func (*Reader) Reset added in v1.3.3

func (z *Reader) Reset(r io.Reader) error

Reset discards the Reader's state and makes it decompress r, keeping its options and decoder, so that a pool of Readers allocates once. It makes a closed Reader usable again, and like NewReader it reads nothing.

type Result

type Result struct {
	Size        int64
	SHA256      [sha256.Size]byte
	ETagEnabled bool
}

Result describes an encoded representation. SHA256 covers exactly Size bytes written to the destination, including the Zstandard frame headers.

func EncodeAll

func EncodeAll(src []byte, options ...Option) ([]byte, Result, error)

EncodeAll encodes src, returning the frame and its cache metadata. The digest is produced while the frame is written; the encoded bytes are not traversed a second time.

Example
package main

import (
	"fmt"

	"github.com/shibukawa/tinygodriver/compress/zstd"
)

func main() {
	encoded, result, err := zstd.EncodeAll([]byte("response body"))
	if err != nil {
		panic(err)
	}
	fmt.Println(len(encoded) == int(result.Size))
	fmt.Println(zstd.ContentEncoding)
}
Output:
true
zstd

func (Result) ETag

func (r Result) ETag() string

ETag returns a quoted strong HTTP entity-tag for the encoded representation. It returns an empty string when the encoder used WithETag(false).

type Writer

type Writer struct {
	// contains filtered or unexported fields
}

Writer emits one Zstandard frame through github.com/klauspost/compress/zstd. Writer is not safe for concurrent use. Close must succeed before Result can be read.

func NewWriter

func NewWriter(w io.Writer, options ...Option) (*Writer, error)

NewWriter starts a host-Go Zstandard stream. Constructing a Writer writes nothing to the destination, so an encoder built and then abandoned leaves the destination untouched. TinyGo and builds using the force_tinygo_logic tag select the bounded TinyGo encoder instead.

func (*Writer) Close

func (z *Writer) Close() error

Close finishes the frame. It does not close the destination.

func (*Writer) Flush added in v1.0.4

func (z *Writer) Flush() error

Flush emits the buffered input as complete blocks so that everything written so far can be decoded, and returns once those bytes reach the destination. It does not end the frame and it does not flush the destination itself. Flushing before a block fills reduces the compression ratio.

func (*Writer) Reset added in v1.2.4

func (z *Writer) Reset(w io.Writer)

Reset starts a new frame writing to w, keeping the encoder and its window so that a pooled Writer does not rebuild them. The ETag setting chosen at NewWriter is retained, nothing reaches w until the caller writes, flushes, or closes, and a nil w leaves the Writer in the error state NewWriter would have reported.

func (*Writer) Result

func (z *Writer) Result() (Result, error)

Result returns the encoded size and SHA-256 digest after a successful Close.

func (*Writer) Write

func (z *Writer) Write(p []byte) (int, error)

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL