video

package
v0.7.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 4, 2026 License: Apache-2.0 Imports: 11 Imported by: 0

Documentation

Overview

Package video is the LOCAL half of codeaf's video work: the things done to files that already exist, on this machine, with ffmpeg — as against the renders bought from a provider a clip at a time (internal/session/tools_video.go).

It exists because the two halves are not the same kind of work and were being treated as one. A render is minutes of somebody else's GPU, costs real money, and produces exactly one short clip because that is what the endpoints do. But a *film* is several of those clips joined, with a frame carried from each one into the next so the shots connect, and a score laid underneath — and every one of those four operations is a local, free, deterministic thing that ffmpeg has done for twenty years.

Until this package, codeaf did them by writing ffmpeg command lines into the shell, and the manual said so in as many words. That has two costs and the second one is the expensive one:

  • It asks a language model to write a filter graph, which is a dialect it knows unevenly and cannot test before running.
  • IT DROPS THE SOUND. The natural way to write a join — `concat` on the video streams, or an `xfade` between them — carries the first input's audio and silently discards every other input's. The result plays, looks right, and goes quiet after the first clip. Nothing errors, nothing warns, and it is not visible in anything but the ear. The manual grew a whole section about the defect ("Why is a stitched video incoherent, or silent after the first clip?") because it kept happening.

So the audio is carried BY CONSTRUCTION here, not by remembering: Join probes every clip, gives the silent ones a generated silence track of their own measured length, and concatenates N video streams with N audio streams. A join cannot come out silent, because there is no code path in which the audio is not mapped.

The library is deliberately free of the session: it takes paths and returns facts and errors, so the belt verb (internal/session/tools_editvideo.go) is a thin wire over it and the same operations are available to anything else that grows a need for them.

Index

Constants

View Source
const Ceiling = 5 * time.Minute

Ceiling is how long one local operation may run before it is given up on, and it is THE number: the belt verb interpolates it into its own description rather than typing a second copy, on this codebase's one-source-of-truth law.

Five minutes is chosen against the work rather than against the clock. A join re-encodes, so it costs roughly the playing time of the material on a modern machine and a handful of ten-second clips is well under a minute; five minutes is a join of many minutes of footage, which is past the point where a person should be watching a tool call spin. What happens at the ceiling matters more than where it is: the half-written file never reaches the destination at all ([produce] encodes elsewhere and renames), so a timeout leaves nothing that could be mistaken for a finished cut and nothing missing that was there before.

View Source
const Closing = time.Duration(-1)

Closing is the moment that means the clip's FINAL frame, for SaveFrame.

It is its own moment rather than a time the caller works out, because the obvious arithmetic is wrong: seeking to the measured length lands one frame past the end and decodes nothing, and seeking to length-minus-a-bit needs the frame rate to know what "a bit" is. ffmpeg can seek relative to the end, so this is the fact and not the estimate.

Variables

View Source
var ErrMissing = errors.New("ffmpeg and ffprobe are not on this machine")

ErrMissing is the one answer every entry point gives on a machine that has no ffmpeg. It is a sentinel rather than a string because the belt uses it for the absence law — a verb with no ffmpeg behind it is left OFF the belt entirely (design-law: a capability that cannot work is absent, not broken), so this error should never reach a model at all.

Functions

func Available

func Available() bool

Available is Missing as the belt asks it.

func Missing

func Missing() string

Missing names the binary this machine has not got, and "" when it has both. It looks the binaries up LIVE rather than caching the answer: the lookup is a few stats against PATH, it is asked once per belt build and never in a loop, and a cached "no" would outlive an ffmpeg installed while codeaf was running.

func SaveFrame

func SaveFrame(ctx context.Context, source string, at time.Duration, destination string) error

SaveFrame writes one frame of source to destination as an image.

at is a time from the start of the clip, or Closing for the final frame. The final frame is the one that matters most: it is how one generated clip connects to the next, handed to generate_video as the opening frame of the shot that follows, and it is the only way codeaf has of making two independent renders look like one continuous take.

Types

type Clip

type Clip struct {
	Path  string
	Facts Facts
}

Clip is one input to a join: where it is and what the probe learned about it.

type Facts

type Facts struct {
	Length time.Duration
	Sound  bool
	Width  int
	Height int
	Rate   float64 // frames per second, 0 when no stream stated one
}

Facts is what one probe learned. Every field obeys the EMPTINESS LAW: a fact the file did not state is the zero value, and a caller reporting it says nothing rather than guessing. That matters most for Length, because the whole point of measuring is that a stitch is timed from it — a length that was guessed would be a cut that drifts.

Sound is a DECLARED audio track, not audible samples: a track of pure silence answers true, exactly as internal/session's mp4 box reader answers it, because the question both are asked is "does the join have an audio stream to carry".

func Join

func Join(ctx context.Context, paths []string, destination string) (Facts, error)

Join concatenates clips into destination and CARRIES EVERY CLIP'S AUDIO.

It probes first, both because the graph is built from the measurements and because a clip that cannot be measured is refused BY NAME before anything is encoded: a join whose third input has no readable duration would otherwise produce a cut that is silently wrong from the third clip on.

The answer is a probe of the RESULT, not a claim about it. What the caller says to the person — how long the cut runs, whether it has sound — is read back off the file that now exists, so it cannot be a promise the encode failed to keep.

func Probe

func Probe(ctx context.Context, path string) (Facts, error)

Probe measures one file. It is ffprobe's json, read for the five facts the rest of this package and the belt verb need, and nothing else.

func Score

func Score(ctx context.Context, clip, audio, destination string, scoring Scoring) (Facts, error)

Score lays an audio file under a video, LOOPED OR TRIMMED to the video's own length — whichever the two lengths call for, without the caller working out which.

That is the whole of what used to be manual arithmetic. generate_music has no length argument (the model writes a piece of its own choosing), so a score almost never matches the cut it is for, and putting one under the other meant measuring both and then looping or trimming by hand. Here the loop is infinite and the output is bounded by the video, so a short piece repeats and a long one stops, and both come out exactly as long as the picture.

The picture is not re-encoded — it is copied stream-for-stream — so scoring a finished cut costs seconds and loses no quality.

type Scoring

type Scoring struct {
	// Level is how loud the score is, 1 being as recorded. There is no default
	// here on purpose: mixing under dialogue and replacing silence want very
	// different numbers, and the caller knows which it is doing.
	Level float64

	// Replace drops the clip's own audio instead of mixing the score under it.
	Replace bool

	// Fade is how long the score takes to fade out at the end. Zero is an
	// honest hard stop; a looped score cut mid-phrase is what a fade is for.
	Fade time.Duration
}

Scoring is how a score sits under a cut.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL