Skip to content

Performance Testing

Use tools/perf/benchmark_effects.py for repeatable effect performance checks. The harness uses only the Python standard library, disables frame-rate sleeps, renders through the library iterator API, and avoids terminal output.

Baseline

Run a baseline before changing performance-sensitive code:

./.venv/bin/python tools/perf/benchmark_effects.py \
  --effect wipe \
  --input-preset medium \
  --json-out /tmp/tte-baseline.json

Use --effect all for a broader pass, and use --input-preset small|medium|large|wide|tall|color|sparse|unicode|generated to select the input shape. Defaults are --samples 7, --warmups 2, and --seed 1337.

The sparse preset creates an 80-by-24 virtual canvas containing only two input characters. It is intended to expose costs that scale with canvas area instead of visible character count; generated provides a dense 80-by-24 control. The unicode preset combines ASCII with CJK, full-width, and single-code-point emoji symbols to exercise double-cell layout and rendering. The wide preset remains a long row of single-cell ASCII and measures input shape, not glyph display width.

Pass --memory to record peak traced Python memory for iterator construction and for the complete iteration. Memory tracking adds substantial runtime overhead, so compare timing results only with another --memory report. Report comparison includes memory deltas automatically when both inputs contain memory measurements.

The default --lifecycle iterator mode measures iterator construction and frame generation without output. Use --lifecycle terminal-output to include the complete public context-manager workflow: terminal preparation, iterator construction, frame generation, in-memory calls to Terminal.print(), and cursor restoration. This mode is useful for finding duplicated setup or output-lifecycle regressions without writing control sequences to the real terminal.

Candidate

After making a focused change and running the relevant functional checks, rerun the same benchmark arguments:

./.venv/bin/python tools/perf/benchmark_effects.py \
  --effect wipe \
  --input-preset medium \
  --lifecycle terminal-output \
  --json-out /tmp/tte-candidate.json

Compare the two reports:

./.venv/bin/python tools/perf/benchmark_effects.py \
  --compare /tmp/tte-baseline.json /tmp/tte-candidate.json

The comparison is advisory by default. Report build, render, and total mean deltas along with frame-count or output-size changes, but do not treat regressions as failures unless a task explicitly sets a threshold.

Profiling

Use --profile when the timing delta needs a call-level explanation:

./.venv/bin/python tools/perf/benchmark_effects.py \
  --effect wipe \
  --input-preset medium \
  --samples 1 \
  --warmups 0 \
  --profile

The profile reports cumulative time for one scenario and is best used after a benchmark shows a meaningful change.

Opt-in row caching

An effect can call self.terminal.enable_row_cache() to reuse formatted rows until their characters change. The engine tracks coordinates, animation visuals, visibility, layers, and direct formatted_symbol assignments. Wide characters use the existing width-aware renderer; caching resumes when all visible characters are single-cell. Existing motion and animation references, paths, and scenes remain valid when caching is enabled.

Measure before enabling caching: tracking and row membership consume extra memory, and frequent changes can reduce the benefit. Matrix enables it for canvases with at least four rows and 256 cells, where fixed-clock measurements showed a benefit. Smaller canvases retain the default renderer. For repeatable Matrix comparisons, advance its wall-clock-based rain phase with the same simulated clock in both runs and verify frame counts and output hashes.

Print, Pour, BouncyBalls, ErrorCorrect, SynthGrid, and Rain also enable caching automatically in measured input ranges:

Effect Activation requirements
Print Text at least 16 columns by 4 rows, with at least 256 input characters.
Pour Up/down pouring; text at least 64 by 12, with at least 768 characters and three-quarter occupancy.
BouncyBalls Text at least 64 by 16, with at least 1,024 characters and three-quarter occupancy.
ErrorCorrect Text at least 8 by 5, with at least 80 input characters.
SynthGrid Visible canvas at least 8 by 4, including sparse text.
Rain Text at least 80 by 24, with at least 1,680 characters and seven-eighths occupancy.

These six effects require single-cell input symbols. The text-based gates also require the text bounds to fit within the visible canvas. Occupancy is the input-character count divided by the area of the text bounds; fill and helper characters do not contribute. These checks run before effect build and helper allocation. Later wide visual changes still use the engine's width-aware fallback. Horizontal Pour and input shapes below these thresholds retain ordinary rendering. Cache activation changes rendering work without changing effect options, scheduling, or random decisions.

Spotlights illumination

Spotlights indexes authored input cells by column and row. Beam selection uses the original floating ellipse spans, including wide continuation cells and input spaces with parsed colors. Empty fill cells are excluded from the index; the effect does not need to materialize them. Input coordinates and symbols remain fixed throughout illumination.

Brightness results use exact color-pair and factor keys in a separate 4,096-entry cache for each iterator. Both initial dim colors and beam falloff use this bounded cache. Factors are not quantized, and cached colors remain immutable.

The effect also reuses an unchanged current visual. Its appearance snapshot includes all visual instance fields and the animation color policy. Direct style or formatted-symbol edits, visual replacement, and policy changes cause the effect to reapply the intended appearance. In always input-color mode, Spotlights uses the shared Animation.set_appearance_if_changed() helper to compare effective colors after input overrides. Other modes retain the local appearance cache. Callers should not rely on receiving a new visual object each frame. The shared set_appearance() method still creates a fresh visual on every call.

Spotlights retains the default terminal renderer. Row caching offered little additional benefit after illumination optimization and regressed the tested Unicode input.

Local paired iterator measurements against the implementation before these changes used seven samples, two warmups, and seed 1337, without terminal printing or frame-rate sleeps. Mean total times (build plus rendering) were:

Input preset Before (s) After (s) Reduction
medium 0.08475 0.04139 51.2%
generated (dense 80-by-24) 2.30201 0.94837 58.8%
sparse 0.29153 0.01922 93.4%
unicode 0.10883 0.05014 53.9%
tall 1.08830 0.37138 65.9%

All paired frame-count and output-length arrays matched. Full-frame hashes also matched across seeded shape, color, and effect-option comparisons. A separate fresh-process tracemalloc sample for dense input at seed 1337 reduced peak Python allocation from 25.46 MB to 8.27 MB (67.5%); build peak increased from 4.43 MB to 5.12 MB. These are local timing and traced-allocation results, not RSS measurements or universal performance guarantees.

Opt-in appearance reuse

Animation.set_appearance_if_changed(symbol=None, colors=None) follows the ordinary setter's symbol defaults, color overrides, and width tracking. It reuses its last visual when the effective request, color policy, visual identity, and all visual instance fields remain unchanged. Scene stepping, an ordinary setter call, or direct style, color-code, width, or formatted-symbol edits force the next helper call to reapply the requested appearance. The helper does not stop an active scene from overwriting that appearance on its next step.

The snapshot is allocated lazily for each animation using the helper. It retains one visual and field snapshot; visuals are not shared across characters. Animation.set_appearance() remains available when callers require a fresh visual on every call. Effects should opt in only after measuring their workloads: comparing and retaining snapshots can cost more than it saves when effective appearances change frequently.

Overflow row coloring and Spotlights opt in only when existing_color_handling="always". For input-derived characters, the helper ignores requested colors in the comparison because parsed input colors override them. Helper characters without input-color ownership still compare the requested colors. Other color modes retain their previous appearance paths.

Local paired iterator benchmarks used seven samples, two warmups, and seeds 1339–1345 after warmup seeds 1337–1338. There was no terminal printing or frame-rate sleeping. Dense always mode results were:

Effect Before (s) After (s) Reduction
Overflow 0.37166 0.25213 32.2%
Spotlights 0.93603 0.77500 17.2%

Frame counts and output lengths matched in all eight timing comparisons. Full-frame hashes matched in 100 seeded comparisons across medium, Unicode, colored, sparse, and dense input, all color-handling modes, and no-color/XTerm rendering. A separate fresh-process dense always sample at seed 1337 changed Overflow peak traced Python allocation from 16.81 MB to 20.43 MB (+21.6%). A separate fresh-process dense always sample at seed 1337 changed Spotlights peak traced Python allocation from 8.07 MB to 8.01 MB (-0.8%). Overflow retains snapshots for moving rows, trading memory for reduced repeated visual construction. These are local timings and traced Python allocations, not RSS measurements or universal speed guarantees.

Motion activation distances

Motion.activate_path() calculates straight-line origin distances with the same terminal-scaled math.hypot arithmetic as the geometry helper, avoiding its coordinate-key hashing. Bézier activation retains the shared curve-distance cache. Every activation still constructs a fresh origin waypoint and segment, resets playback, and dispatches events in the same order. No per-path distance cache is allocated.

Local production verification on 2026-10-03 used seven paired samples, two warmups, and measured seeds 1339–1345, without terminal printing or frame-rate sleeps. The baseline method was snapshotted before editing; all other engine and effect code was shared. Geometry caches were cleared before each iteration, then warmed by its own build.

Rings input Before total (s) After total (s) Reduction
generated 1.41773 1.38221 2.5%
medium 0.10063 0.09800 2.6%
tall 0.42941 0.41499 3.4%
wide 0.01797 0.01755 2.3%
sparse 0.00674 0.00657 2.4%
unicode 0.08290 0.08093 2.4%

All per-seed frame-count and output-length arrays matched, as did 44 seeded full-frame comparisons across Rings shapes, colors, offsets and configurations, and six control effects. Separate dense seed-1337 traced allocation samples measured 110.56 MB before and 111.67 MB after (+1.0%). These are local iterator timings and traced Python allocations, not RSS measurements or universal speed guarantees.

Opt-in eased-scene schedules

Animation.new_scene(cache_easing=True) or Scene(..., cache_easing=True) reuses immutable frame-index schedules for scenes with matching built-in easing functions and duration boundaries. Frames, visuals, playback counters and lifecycle events remain independently owned. Only Waves opts in by default. Custom easing callables, scenes beyond the 4096-tick entry limit, and motion-synced scenes retain their ordinary playback calculations. The engine keeps at most 32 schedules; weak scene references allow entries to be released on eviction and lazily reacquired during playback. Easing changes or appended frames select a new schedule; resetting/looping reuses the existing layout.

Use the helper after measuring repeated layouts. Many unique schedules or short/cancelled playback can pay for precomputation without enough reuse. Opting in does not cache arbitrary callback results or share mutable animation state. Manually changed out-of-range cursors retain ordinary easing/clamping behavior.

Local production comparisons used the repository harness with actual pre-change source in isolated workers: seven samples, two warmups, seeds 1339–1345, serial rotating arms, GC before each run and cold schedule caches. Defaults, no terminal printing and no frame-rate sleeping were retained. The baseline already contains Waves' effect-local construction optimizations. Results measure this engine feature alone.

Input / color mode Before total s After total s Iteration reduction Total reduction
generated / ignore 1.849471 1.644767 19.7% 11.1%
medium / ignore 0.068867 0.061437 18.5% 10.8%
unicode / ignore 0.042426 0.038307 15.9% 9.7%
generated / always 1.733739 1.529323 19.4% 11.8%
generated / dynamic 1.630213 1.423876 22.5% 12.7%
sparse / dynamic 0.002995 0.002934 3.6% 2.1%

All timing and separate memory frame/output-count arrays matched. All 68 full-frame SHA256/final-state comparisons matched across Waves shape/color/easing/configuration cases and BinaryPath/Expand controls. A fresh-process dense seed-1339 trace changed peak Python allocation from 154.208 MiB to 154.402 MiB (+0.126%, about 0.19 MiB). Memory tracing ran outside timing. BinaryPath/Expand kept caching disabled; their small total-time variations (+0.9%/-1.8%) did not establish a meaningful change. Thunderstorm remains disabled and has no timing claim because its output is not reproducible under the existing benchmark. These are local measurements, not universal guarantees or RSS.

Opt-in encoded appearance construction

Animation.new_scene(cache_appearance=True) or Scene(..., cache_appearance=True) shares immutable ANSI strings during frame construction. Every frame and visual remains a fresh mutable object. The engine retains at most 2048 encodings, keyed by symbol, style mask and colors resolved after terminal policy. Only native inputs use the cache; custom values and formatting overrides preserve ordinary behavior. Explicit formatting calls remain uncached, and visual edits keep their existing output invalidation.

The option defaults to disabled. Waves enables it for the repeated wave scene. Use it after measuring both construction time and peak allocation: a high hit rate alone does not establish a speedup.

Production measurements against the already optimized Waves with easing schedules, using seven samples, two warmups, rotated isolated workers and cold caches per iteration:

Dense Waves mode Build delta Iteration delta Total delta Peak allocation delta
ignore +1.2% -3.6% -1.3% -18.1%
always +6.8% -2.7% +1.6% -10.0%
dynamic -0.1% -5.3% -2.6% -19.6%

The main benefit is allocation reduction: dense ignore peak decreased from 154.4 MiB to 126.5 MiB. Medium/Unicode/always totals measured 1.6–2.8% slower. Disabled BinaryPath/Expand controls measured about 1% slower in the final run and slightly faster in the initial run; no control speedup is claimed. Construction metadata and guards still have a cost, so further enablement requires separate measurements. Timing frame/output counts matched, as did 68 paired complete-animation hashes and final states. Memory tracing and correctness checks ran outside timing.