Skip to content

0.0.16 plan — data and developer tooling

Status: proposed on 2026-10-06, after the 0.0.15 cohort was published. The scope, and the decisions marked "decided", were agreed with the maintainer the same day. It tracks the first slice of milestone 6 (Data & developer tooling), plus ER diagrams (#247) from milestone 5 and DOT cluster frames from the 0.0.15 known limits. The scope tracker is #660.

Goal: structured data you can read and trust in a terminal.

  • Tables from real data. One CSV/JSONL/serde adapter layer, column types inferred on request, column statistics, conditional styles, and tables too large to hold on screen, so a file or a query result becomes a table without hand-written glue.
  • Profiles, schemas and records. rich profile for a CSV or JSONL file; schemas from JSON Schema, Arrow and SQL compared in one model and drawn as ER diagrams; a record inspector that drills into nested data.
  • Developer views. Three-way merge conflicts, and Cargo feature, size, build-time, advisory and licence reports beside rich deps.

What stays the same: core's behaviour. A default build still behaves like upstream rich 15.0.0, and nothing under crates/rich/src changes. rich --csv keeps rich-cli 1.8.1's output byte for byte.

Research behind this plan

A read-only survey on 2026-10-06:

  • Much of the base exists. rich_ext::data has the Node model, JSON/INI/dotenv with YAML/TOML/XML behind features, Explorer, TableView, flatten, search and redaction. rich_ext::table has a typed Value, Column, sorting, GroupBy with aggregates and StreamingTable. rich_ext::schema has SchemaTree and SchemaDiff (JSON Schema only); rich_ext::deps has DepGraph, duplicates() and --why; rich_ext::diff has the Myers engine, SourceDiff and test_report; rich_ext::chart has Histogram and Heatmap.
  • There is no tabular input layer. CSV is read by one CLI-private reader, read_csv_rows in crates/rich-cli/src/lib.rs, which --csv uses and rich chart's parse_csv (crates/rich-cli/src/chart.rs) wraps; nothing outside the CLI can use it. JSONL is not read anywhere. No crate depends on csv, arrow, polars, parquet or a SQL parser.
  • The milestone's issues are one or two sentences each, so this plan sets the scope. None needs a core change.
  • The 0.0.15 known limits that matter here: DOT clusters are drawn without frames and rank is ignored (ER diagrams and lineage want clusters), and schema trees stop at 10,000 entries.

Workstreams

Each workstream lands as its own pull request, or a short stack, with evidence in the PR and an entry under Delivered. Placement follows AGENTS.md: nothing in core; general renderables in rich-ext; tabular adapters, profiling and heavy formats in a new rs-rich-data crate; graph drawing in rs-rich-diagram; the CLI only composes.

1. Tabular foundation — data (new), ext, CLI

  • The minimal schema model first: the format-neutral field model that workstream 4 builds on (fields with a type, nullability and children) lands in rich_ext::schema as this workstream's first step, so the row source can carry a schema without a temporary type.
  • A new crate, rs-rich-data (rich_data), with a row source contract (column names, an optional schema in that model, rows of table::Value) and adapters behind it: CSV and TSV (the CLI's reader, moved and shared by --csv, rich chart and everything below), JSONL, and serde (Serialize rows). Arrow RecordBatch sits behind an off-by-default arrow feature. Tabular adapters for serde, CSV, Arrow and Polars (#214): Polars is deferred.
  • Type inference (#267): integers, floats, booleans, dates and timestamps, nulls and text per column, with the evidence (how many values parsed) and an override. Opt-in: --csv keeps upstream's complex() justification and never infers.
  • Column statistics (#261): count, nulls, distinct, min, max, mean, median and quantiles, and the top values, as a table or under each column heading.
  • Conditional styles (#215): rules that style a cell, row or column by value (column, a comparison, a value, a style), from Rust predicates or TOML rule tables in the existing config profiles. No expression language.

Acceptance: rich --csv output and its tests are byte-identical after the reader moves; each adapter reads its fixtures to the same rows; inference is reported with its evidence and never changes --csv.

2. Large and SQL-shaped tables — ext, data

  • Virtualised tables (#260): a renderer that slices rows to a viewport with column widths fixed or sampled from the first N rows, so a million-row source renders in constant memory. Interactive scrolling reuses the rs-rich-interact viewport; the milestone-4 table extras (#430–#433) stay there.
  • SQL result renderer (#235): a result set with its column types (from a row source's schema) shown with typed alignment, NULL styling and the row count. No database connection: rows come from an adapter or the caller.

3. Profiling and data quality — data, CLI

  • rich profile FILE: a CSV (#343) or JSONL (#344) profile with per-column type, null rate, distinct count, statistics and a distribution (#345: numeric histograms from chart::Histogram, top-N bars for categories), and a missing-value map (#346) built on chart::Heatmap. --report json for machines.
  • Data quality results (#269): a generic check-result model (check, column, status, observed, expected) rendered like test_report, with a summary and the failing rows.

Acceptance: profiles of the fixture files match their snapshots; files larger than memory are sampled, and the sample size is printed.

4. Schemas — ext, data, diagram

  • The format-neutral schema model in rich_ext::schema, extended from the minimal model workstream 1 adds: constraints, and the mappings from JSON Schema, Arrow and SQL DDL. The existing SchemaTree and SchemaDiff are rebuilt on it with byte-identical output for JSON Schema.
  • Schema diff (#268): added, removed and changed fields with breaking changes marked, for any two schemas the model can read.
  • Arrow schema explorer (#342): behind rs-rich-data's arrow feature, as a tree and through the diff.
  • Schema evolution timeline (#347): a series of schema versions on chart::Timeline, with each change marked.
  • ER diagrams (#247): tables, columns and keys from SQL DDL (a small CREATE TABLE subset) or the schema model, drawn through rs-rich-diagram.
  • DOT cluster frames (#240): the layout draws cluster frames, so ER groups and DOT subgraph cluster_* show their boxes; rank=same is honoured where the layered layout can.

5. Records and conflicts — ext

  • Record inspector (#270): one record shown as a table of fields with nested values folded into drill-down trees, built on data::Node and Explorer, with a hook the interactive explorer uses to open a branch.
  • Three-way merge conflicts (#339): a parser for conflict markers (including diff3's base section) and a view with ours, base and theirs side by side or stacked, syntax highlighted, as rich diff --conflicts FILE. The interactive resolver (#467) stays in milestone 4 and builds on it.

6. Cargo and the supply chain — ext, CLI

All of these read files or JSON that Cargo and its tools already write; none run a scanner or reach the network, following deps.rs's rule that running Cargo is the caller's business.

  • Duplicates (#420): a consolidation summary on top of DepGraph::duplicates (rich deps --duplicates already lists them).
  • Feature trees (#419): cargo metadata's resolved features per crate, as rich deps --features.
  • Sizes and build times (#421): cargo build --timings output as bars and a table, as rich deps --timings FILE.
  • Advisories (#330, #418): a generic advisory model, read from cargo audit --json, as rich deps --audit FILE. #418 asks for a plugin; the renderer ships built in instead, and a plugin can reuse it.
  • Licences (#331): licences from cargo metadata, grouped, with unknown and copyleft licences marked, as rich deps --licenses.

7. The CLI, Python and docs — CLI, py, docs

  • CLI: rich profile; rich schema reading Arrow, SQL DDL and the ER view (--er); rich deps --features/--timings/--audit/--licenses; rich diff --conflicts; a --infer option where tables come from files.
  • Python: rs_rich.data bindings for the adapters, profiles and tables, and the new schema and conflict views.
  • Docs: a data guide (adapters, inference, statistics, styles, profiles), a schema guide, a developer-views guide, tapes for rich profile, rich schema --er, rich deps --features and rich diff --conflicts, and gallery entries.

Order of work

  1. Workstream 1 first, starting with the minimal schema model: everything else reads rows through it, and workstream 4 extends that model.
  2. Workstreams 5 and 6 in parallel with 1; they share nothing with it except table::Value.
  3. Workstream 4 once workstream 1's schema model has merged.
  4. Workstreams 2 and 3 after 1's row source and inference settle.
  5. Workstream 7 last, since it shows everything else.
  6. Release test.

Decisions

Decided with the maintainer on 2026-10-06 unless marked as a default.

  • Scope (decided): milestone 6's tabular, profiling, schema, record, conflict and Cargo issues, with ER diagrams (#247) and DOT cluster frames. The other milestone-6 themes wait (see below).
  • A new crate, rs-rich-data (rich_data) (decided), for the tabular adapters, inference, statistics, profiling and the heavy formats, so rich-ext stays light. This is a deliberate exception to AGENTS.md's rule that new library-facing features go in crates/rich-ext, of the same kind as rs-rich-diagram, rs-rich-micro and rs-rich-interact: a crate with its own reason to exist (heavy optional dependencies such as Arrow, and formats ext's users should not compile), which only builds on ext's public API and never touches core. General renderables it needs (conditional styles, virtualised tables, the schema model) still go in rich-ext. It depends on rich-ext and rs-rich-diagram (for ER diagrams); rich-cli and rich-py depend on it. Arrow sits behind an off-by-default arrow feature; Polars is deferred; Parquet, Delta Lake and dbt come later as plugin crates. It starts at 0.0.1, with a first upload through the release workflow's new-crate path, and the pull request that creates it adds it to the dependency graph and the versioning table in AGENTS.md.
  • One shared CSV reader, --csv unchanged (decided): the CLI's hand-written reader moves into rs-rich-data with no new dependency; rich --csv keeps rich-cli 1.8.1's behaviour, and inference is opt-in for the new commands.
  • A small CLI surface (decided): one new command, rich profile; the rest extends rich schema, rich deps and rich diff rather than a command per issue.
  • Conditional styles are rules, not a language (default): Rust predicates and TOML rule tables in the existing config profiles.
  • Virtualisation is a static renderer (default) in ext; interactive paging reuses rs-rich-interact's viewport.
  • Supply-chain reports read existing output (default): cargo metadata, cargo build --timings and cargo audit --json; nothing runs a scanner or reaches the network.
  • No core change (default). If something needs a core hook, it goes through an extension-point trait in protocol.rs, recorded in DIVERGENCES, and rs-rich moves to 0.0.10. None is expected.
  • If the release has to shrink (default), drop in this order: the SQL result renderer (#235), the schema timeline (#347), licences (#331), ER diagrams (#247), then virtualised tables (#260). The priority:high issues (#214, #267, #268, #270, #339) stay.

Not in this release

  • Milestone 6, later: query plans and lineage (#348, #349, and the Spark, database and dbt plugins #350, #351, #352, #354), the SQL formatter (#353), Parquet and Delta Lake (#340, #341), Git history, blame, changelog and source outline views (#336, #337, #338, #327, #328), test history, coverage and flamegraphs (#333, #334, #335), and SBOMs (#332).
  • Milestone 5, later: OpenAPI (#245), AsyncAPI (#246), PlantUML, D2 and Vega (#241, #242, #243), grouped progress dashboards (#255), treemaps (#253) and call graphs (#249).
  • Polars and any database connectivity.
  • Anything that makes core aware of data formats. Core stays a faithful mirror.

Package impact

Package Proposed Reason
rs-rich 0.0.9 (unchanged) No core change is planned
rs-rich-data 0.0.1 (new) Adapters, inference, statistics, profiling, Arrow
rs-rich-ext 0.0.14 Schema model, conditional styles, virtualised tables, record inspector, conflicts, supply-chain views
rs-rich-diagram 0.0.2 Cluster frames, rank, ER diagrams
rs-rich-mermaid 0.0.5 Manifest only: requires diagram 0.0.2
rs-rich-micro, rs-rich-interact, rs-rich-record next patch Manifest only: require ext 0.0.14
rs-rich-cli 0.0.16 rich profile, and the schema, deps and diff options
rs-rich (PyPI) 0.0.5 rs_rich.data and the new views
rs-rich-art, rs-rich-plugin-api, rs-rich-lumis, rs-rich-macros unchanged Nothing they need changes

Cargo reads a 0.0.x requirement as an exact version, so every crate that depends on a bumped crate moves with it. RELEASES.toml records the order, and python scripts/release_cohort.py tag publishes it.

Delivered

Workstream Pull request Merged as
Plan #661 c3e69f0
1. Tabular foundation #662 e027731
5. Records and conflicts #663 31c532c
6. Cargo and the supply chain this PR open

Review focus

  1. Core is untouched. git diff shows nothing under crates/rich/src, and every golden and the differential corpus are byte-identical.
  2. rich --csv did not move. Its tests and the rich-cli parity cases are byte-identical after the reader moves into rs-rich-data.
  3. JSON Schema output did not move. rich schema snapshots are byte-identical after SchemaTree and SchemaDiff are rebuilt on the format-neutral model.
  4. Large inputs stay bounded. Profiles sample, virtualised tables hold one viewport, and every parser has a size or depth limit like rs-rich-diagram's.