Skip to content

Quality assurance

rich_ext::qa tests how your output renders, not only what it says. It renders a renderable through explicit capabilities, never the real terminal, so results are the same on every machine and in CI.

Module Answers
screenshot Did the output change? Approved text files, with a review workflow
stress Does it survive every width and height?
lint Is anything clipped, mis-styled, colour-only or unsupported by the target?
explain Why did it wrap, truncate, lose colour or fall back to ASCII?
profile How long do measure, render and a live refresh take, and how much do they allocate?
fuzz Do random renderables break any invariant?
matrix Does it hold across 16 terminal profiles?
bench Did it get slower than the baseline?

Everything here needs the testing feature. Add it as a dev-dependency so it never reaches your release build:

[dev-dependencies]
rs-rich-ext = { version = "0.0.9", features = ["testing"] }

The examples come from guide_qa.rs:

cargo run -p rs-rich-ext --example guide_qa --features testing

They all use this fixture:

fn status_table() -> Table {
    let mut table = Table::new().title("Deployments");
    table.add_column("Service");
    table.add_column("Region");
    table.add_column("Status");
    let status = |markup: &str| Text::from_markup(markup).expect("valid markup");
    table.add_row_text(vec![
        Text::new("api"),
        Text::new("eu-west-1"),
        status("[green]✔ ok[/]"),
    ]);
    table.add_row_text(vec![
        Text::new("worker"),
        Text::new("us-east-2"),
        status("[red]✖ failed[/]"),
    ]);
    table
}

Wire it into cargo test

Each tool returns a report you can assert on, so a render test is an ordinary #[test]:

#[cfg(test)]
mod tests {
    use super::*;
    use rich_ext::qa::fuzz::NodeKind;
    use rich_ext::qa::screenshot::assert_screenshots;

    #[test]
    fn status_table_screenshots() {
        // Approved files live in tests/screenshots/status-table/; run with
        // RICH_APPROVE=1 to accept new output.
        assert_screenshots("status-table", &status_table());
    }

    #[test]
    fn status_table_survives_common_widths() {
        let report = stress(&status_table(), &StressOptions::widths([20, 40, 80, 120]));
        assert!(
            report.is_clean(),
            "{}",
            Console::new().render_to_string(&report)
        );
    }

    #[test]
    fn status_table_lints_clean() {
        let findings = lint(&status_table(), &LintOptions::default());
        let report = LintReport::new(findings);
        assert!(!report.has_errors(), "{}", report.to_json());
    }

    #[test]
    fn fixtures_hold_across_unicode_terminals() {
        // The status symbols need Unicode: ASCII profiles fail.
        let profiles: Vec<_> = CapabilityProfile::standard()
            .into_iter()
            .filter(|p| p.unicode())
            .collect();
        let report = matrix::regression(FIXTURES, &profiles, 40);
        assert!(report.is_ok());
    }

    #[test]
    fn fuzzing_finds_nothing() {
        let options = GenOptions::default().kinds([NodeKind::Text, NodeKind::Table]);
        let report = fuzz_with(2024, 300, &options, &Invariants::default());
        assert!(report.is_clean(), "{:#?}", report.failures);
    }
}

The sections below explain each tool.

Screenshot approvals

Screenshot::capture(name, &renderable, &matrix) renders once per combination of a Matrix: widths × colour depths × Unicode on or off. The default matrix is 40, 80 and 120 columns × truecolor and no colour × Unicode and ASCII, which gives 12 shots. Each Shot has a key such as status-table@80.truecolor.unicode.

Approvals::new(dir) stores approved shots as <dir>/<name>/<key>.txt:

  • no-colour shots are the plain text;
  • colour shots are the ANSI output with escapes written visibly (\e[1m), so the files are printable and diff line by line.

check(&shots) compares shots with those files. A shot that differs, or has no approved file yet, is written beside it as <key>.new. The workflow:

  1. Run the tests. New or changed output fails, and .new files appear.
  2. Review the .new files (or the diff in the failure message).
  3. Rerun with RICH_APPROVE=1 to accept them. approve_all() does the same from code, and approving(true) accepts during a check.
  4. Commit the .txt files. A later match removes stale .new files.
fn approval_workflow(dir: &Path) -> String {
    // Two widths, colour and no colour: four shots.
    let matrix = Matrix::default()
        .widths([30, 60])
        .color([ColorDepth::TrueColor, ColorDepth::None])
        .unicode([true]);
    let approvals = Approvals::new(dir);

    // First run: nothing approved yet, so every shot is "missing" and
    // written as `<key>.new` for review.
    let shots = Screenshot::capture("status-table", &status_table(), &matrix);
    let first = approvals.check(&shots).expect("readable directory");
    assert_eq!(first.missing.len(), 4);

    // Accept them (what RICH_APPROVE=1 does during a check).
    approvals.approve_all().expect("writable directory");

    // The renderable changes: the approved files no longer match.
    let mut changed = status_table();
    changed.add_row(&["cron", "eu-west-1", "paused"]);
    let shots = Screenshot::capture("status-table", &changed, &matrix.clone().widths([60]));
    let outcome = approvals.check(&shots).expect("readable directory");
    assert!(!outcome.is_ok());
    outcome.report() // a summary, then a diff per mismatch
}

A failing check: a summary and a diff per mismatched shot

assert_screenshots(name, &renderable) does all of this with the default matrix. It keeps files in RICH_SCREENSHOT_DIR, else $CARGO_MANIFEST_DIR/tests/screenshots, and panics with the diffs. assert_screenshots_with(&approvals, name, &renderable, &matrix) takes your own directory and matrix.

Stress

stress(&renderable, &options) renders at many widths and heights and reports:

  • overflow: a line wider than the width;
  • clipping: content lost compared with a wide render;
  • unstable wrapping: more lines at a larger width;
  • panics;
  • measure mismatch: output that disagrees with measure().

StressOptions::default() uses widths 1, 2, 3, 4, 10, 20, 40, 80, 120 and 200, with the height unset, 5 and 24. StressOptions::widths([…]) picks widths with the height unset.

fn show_stress(console: &Console) {
    let report = stress(&status_table(), &StressOptions::widths([3, 8, 20, 40, 80]));
    if !report.is_clean() {
        console.print(&report);
    }
}

A table that overflows at 3 columns

report.is_clean() is the usual assertion. report.of(IssueKind::Overflow) filters the issues by kind.

Lint

lint(&renderable, &LintOptions) renders for a described target and checks:

  • layout: stress at the lint widths (default 20, 40 and 80);
  • hyperlinks: empty, malformed or unsafe URLs (javascript:), unknown schemes, and links with blank text;
  • colour-only distinctions: the same status symbol or word (●, ✔, ok, error, …) in two colours with nothing else to tell them apart;
  • capabilities: colours deeper than the target, non-ASCII glyphs on an ASCII target, links on a target without OSC 8, and blink.

A theme name that does not exist renders as nothing, so it cannot be found after rendering. lint_markup(markup, &theme) and lint_text(&text, &theme) check the source instead and suggest the nearest name.

fn show_lint(console: &Console) {
    // Markup is checked against a theme: typos get the nearest name.
    let theme = rich_ext::extended_theme();
    let mut findings = lint_markup("[bold gren]ready[/] [repr.nubmer]42[/]", &theme);

    // A render is checked at each width for a described target.
    let options = LintOptions::default().widths([12, 40]).unicode(false);
    let status =
        Text::from_markup("[green]● ok[/] [red]● failed[/] [link=htps://example.com]docs[/]")
            .unwrap();
    findings.extend(lint(&status, &options));

    let report = LintReport::new(findings);
    let _json = report.to_json(); // for CI annotations
    console.print(&report);
}

Lint findings with severities, rules and locations

LintOptions::default() describes a 16-colour Unicode terminal without links. LintOptions::capable() describes truecolor with links. Adjust with widths, color, unicode, hyperlinks and layout. LintReport counts findings by severity, has has_errors(), and round-trips through to_json() / from_json() for CI annotations.

Explain

explain(&renderable, &target) answers "why does it look like that?" for a RenderTarget. It reports which lines wrapped and from where, truncation and cropping, each colour downgrade with its mapping (#ff8700 → 208 → bright_red), the fidelity level and why, ASCII substitutions, and dropped links. explain_console explains for an existing console, and explain_with_report adds each capability's source from a capability report.

fn show_explain(console: &Console) {
    let mut table = Table::new();
    table.add_column("Step");
    table.add_column("Result");
    table.add_row_text(vec![
        Text::from_markup("[#ff8700]deploy[/]").unwrap(),
        Text::new("finished in 42s with no errors"),
    ]);
    // A 16-colour, ASCII-only terminal, 24 columns wide.
    let target = RenderTarget::new(
        TargetKind::Terminal,
        TargetCapabilities {
            width: 24,
            height: 24,
            color_system: Some(ColorSystem::Standard),
            interactive: true,
            unicode: false,
            hyperlinks: false,
            sixel: Support::Unsupported,
        },
        Theme::default_theme(),
    );
    let explanation = explain(&table, &target);
    console.print(&ExplanationView::new(&explanation));
}

Why a table wrapped, lost colour and fell back to ASCII

Profile

profile(&renderable, &console, &options) times measure() and rich_render() over iterations runs (default 20, after 2 warm-up runs). It counts segments, lines, cells and ANSI bytes. With frame(width, height) it also times a live-style refresh: render at a fixed size, shape to exactly that many rows and encode, as Live does on every frame.

To count allocations, install CountingAllocator as the global allocator of the test or bench binary. The library never installs it:

use rich_ext::qa::profile::CountingAllocator;

// In a test or bench binary, never in a library.
#[global_allocator]
static ALLOC: CountingAllocator = CountingAllocator::system();
fn run_profile(console: &Console) -> Profile {
    let options = ProfileOptions::default()
        .iterations(50)
        .width(60)
        .frame(60, 10); // also cost a Live-style 60x10 refresh
    let profile = profile(&status_table(), console, &options);
    // Filled in because this binary installs CountingAllocator.
    assert!(profile.allocations.is_some());
    profile
}

fn show_profile(console: &Console, profile: &Profile) {
    console.print(&ProfileReport::new(profile));
}

A profile report (numbers in this screenshot are illustrative)

The counters are process-wide. Parallel tests allocate at the same time, so treat the numbers as an upper bound, or run with --test-threads=1. The numbers in the screenshot are made up, so this page does not change with the machine that built it.

Fuzz

fuzz(seed, cases, &invariants) generates random renderable trees from a seed: styled text with wide, combining and zero-width characters, tables, panels, padding, alignment, columns and trees. It renders each tree at a random width and checks the Invariants:

  • no panic;
  • no line wider than the width;
  • a second render is identical;
  • measure() bounds hold;
  • your own checks, added with custom(name, check).

A failure is shrunk while the same invariant still fails. It keeps its seed and case index, so Case::generate(seed, index, &options) rebuilds it exactly, and failure.minimized is a Rust reproduction to paste into a test. fuzz_with takes GenOptions: widths, depth, children, words, a shrink budget, and which node kinds to generate (kinds, without).

fn show_fuzz(console: &Console) {
    // Every node kind, widths 4 to 40.
    let options = GenOptions {
        min_width: 4,
        max_width: 40,
        ..GenOptions::default()
    };
    // The default invariants, plus one of our own.
    let invariants = Invariants::default().custom("no-tabs", |r: &Rendered<'_>| {
        match r.lines.iter().position(|line| line.contains('\t')) {
            Some(i) => Err(format!("line {} contains a tab", i + 1)),
            None => Ok(()),
        }
    });
    let report = fuzz_with(7, 60, &options, &invariants);
    console.print(&report);
    for failure in &report.failures {
        // Rebuild the exact case from its seed and index...
        let case = Case::generate(failure.seed, failure.case_index, &options);
        assert_eq!(case, failure.case);
        // ...or paste the shrunk reproduction into a test.
        println!("{}", failure.minimized.as_deref().unwrap_or(""));
    }
}

No failures across 60 generated cases

A failure lists the case, the width, the broken invariant and a shrunk reproduction you can paste into a test, like this one:

// seed 7, case 12, width 1
let renderable = Columns::new(vec!["su".to_string()]);

That reproduction is from a real bug the fuzzer found: Columns used to overflow with an item wider than the width, and Tree guides overflowed below about 8 columns. Both are fixed in core, matching rich 15.0.0.

Capability matrix

matrix::regression(fixtures, &profiles, width) renders each fixture under each CapabilityProfile and checks the structure: nothing wider than the width, no colour codes on a no-colour profile, no OSC 8 links on a no-link profile, and no non-ASCII on an ASCII profile. CapabilityProfile::standard() is 16 profiles: 16, 256 and truecolor × Unicode or ASCII × links or not, then dumb, ci, windows-terminal and screen-reader.

A fixture is a name and a fn() -> Box<dyn Renderable>:

fn fixture_table() -> Box<dyn Renderable> {
    Box::new(status_table())
}

fn fixture_panel() -> Box<dyn Renderable> {
    Box::new(Panel::new(Box::new(Text::styled("all systems go", "green"))).title("Status"))
}

const FIXTURES: &[Fixture] = &[("table", fixture_table), ("panel", fixture_panel)];

fn show_matrix(console: &Console) {
    // 12 colour × unicode × link combinations, then dumb, ci,
    // windows-terminal and screen-reader.
    let profiles = CapabilityProfile::standard();
    let report = matrix::regression(FIXTURES, &profiles, 40);
    console.print(&report);
    for cell in report.failures() {
        eprintln!("{} under {}: {:?}", cell.fixture, cell.profile, cell.status);
    }
}

The status table fails every ASCII profile

The status table fails on ASCII profiles because ✔ and ✖ are not ASCII. Pick status symbols per profile with an AccessibilityPolicy, or wrap the output in Degrade.

matrix::run(…, Some(&approvals)) also checks every cell against approved screenshots.

Benchmarks

Bench::new(name).run(|| …) samples a closure until a target time (default 1 s, at least 10 samples) and returns a Measurement with the mean, median, standard deviation, p95, min and max. bench_renderable(name, &renderable, width) benchmarks a render plus ANSI encoding. A BenchRun holds the measurements and saves to a stable JSON format (save, load). It can also read criterion output with from_criterion_dir.

compare(&baseline, &candidate, &CompareOptions) gives each benchmark a verdict: regression, improvement, unchanged, new or removed. A change must exceed threshold_pct (default 5) and, by default, the combined standard deviation. ComparisonView renders it with a min/median/p95/max sparkline per side.

fn record_bench(path: &Path) -> std::io::Result<()> {
    let table = status_table();
    let run = BenchRun::new(vec![
        bench_renderable("status-table@80", &table, 80),
        bench_renderable("status-table@40", &table, 40),
    ]);
    run.save(path) // JSON; `rich bench compare` reads it
}

fn show_comparison(console: &Console, baseline: &BenchRun, candidate: &BenchRun) -> bool {
    let options = CompareOptions {
        threshold_pct: 10.0,
        ..CompareOptions::default()
    };
    let comparison = compare(baseline, candidate, &options);
    console.print(&ComparisonView::new(&comparison));
    comparison.has_regressions()
}

A comparison with a regression, an improvement, a new and a removed benchmark (illustrative numbers)

From the command line, rich bench compare BASE CAND [--threshold PCT] prints the same table and exits with status 5 when anything regressed, which makes it a CI gate:

rich bench compare baseline.json candidate.json --threshold 10
rich bench compare target/criterion-main target/criterion

Gotchas

  • Determinism. Every tool renders through capabilities you pass, never the environment. Fix the widths and profiles in tests, and do not assert on timings.
  • Approval files for colour shots are escaped ANSI. Review them as text: \e[1;31mred\e[0m is the escape it looks like.
  • RICH_ASSERT_COLOR controls colour in approval diffs, as it does for the assertion macros.
  • Natural width. explain takes the natural width from measure() at a very wide probe. Core Table and Panel do not implement measure() yet, so for them the summary line reports the probe's width (1000), as in the screenshot above. The wrap events themselves count the real content width.

See also