0.0.16 plan — data and developer tooling¶
Status: proposed on 2026-10-06, after the 0.0.15 cohort was published. The scope, and the decisions marked "decided", were agreed with the maintainer the same day. It tracks the first slice of milestone 6 (Data & developer tooling), plus ER diagrams (#247) from milestone 5 and DOT cluster frames from the 0.0.15 known limits. The scope tracker is #660.
Goal: structured data you can read and trust in a terminal.
- Tables from real data. One CSV/JSONL/serde adapter layer, column types inferred on request, column statistics, conditional styles, and tables too large to hold on screen, so a file or a query result becomes a table without hand-written glue.
- Profiles, schemas and records.
rich profilefor a CSV or JSONL file; schemas from JSON Schema, Arrow and SQL compared in one model and drawn as ER diagrams; a record inspector that drills into nested data. - Developer views. Three-way merge conflicts, and Cargo feature,
size, build-time, advisory and licence reports beside
rich deps.
What stays the same: core's behaviour. A default build still behaves
like upstream rich 15.0.0, and nothing under crates/rich/src changes.
rich --csv keeps rich-cli 1.8.1's output byte for byte.
Research behind this plan¶
A read-only survey on 2026-10-06:
- Much of the base exists.
rich_ext::datahas theNodemodel, JSON/INI/dotenv with YAML/TOML/XML behind features,Explorer,TableView, flatten, search and redaction.rich_ext::tablehas a typedValue,Column, sorting,GroupBywith aggregates andStreamingTable.rich_ext::schemahasSchemaTreeandSchemaDiff(JSON Schema only);rich_ext::depshasDepGraph,duplicates()and--why;rich_ext::diffhas the Myers engine,SourceDiffandtest_report;rich_ext::charthasHistogramandHeatmap. - There is no tabular input layer. CSV is read by one CLI-private
reader,
read_csv_rowsincrates/rich-cli/src/lib.rs, which--csvuses andrich chart'sparse_csv(crates/rich-cli/src/chart.rs) wraps; nothing outside the CLI can use it. JSONL is not read anywhere. No crate depends oncsv,arrow,polars,parquetor a SQL parser. - The milestone's issues are one or two sentences each, so this plan sets the scope. None needs a core change.
- The 0.0.15 known limits that matter here: DOT clusters are drawn
without frames and
rankis ignored (ER diagrams and lineage want clusters), and schema trees stop at 10,000 entries.
Workstreams¶
Each workstream lands as its own pull request, or a short stack, with
evidence in the PR and an entry under Delivered. Placement
follows AGENTS.md: nothing in core; general renderables in rich-ext;
tabular adapters, profiling and heavy formats in a new rs-rich-data
crate; graph drawing in rs-rich-diagram; the CLI only composes.
1. Tabular foundation — data (new), ext, CLI¶
- The minimal schema model first: the format-neutral field model
that workstream 4 builds on (fields with a type, nullability and
children) lands in
rich_ext::schemaas this workstream's first step, so the row source can carry a schema without a temporary type. - A new crate,
rs-rich-data(rich_data), with a row source contract (column names, an optional schema in that model, rows oftable::Value) and adapters behind it: CSV and TSV (the CLI's reader, moved and shared by--csv,rich chartand everything below), JSONL, and serde (Serializerows). ArrowRecordBatchsits behind an off-by-defaultarrowfeature. Tabular adapters for serde, CSV, Arrow and Polars (#214): Polars is deferred. - Type inference (#267): integers, floats, booleans, dates and
timestamps, nulls and text per column, with the evidence (how many
values parsed) and an override. Opt-in:
--csvkeeps upstream'scomplex()justification and never infers. - Column statistics (#261): count, nulls, distinct, min, max, mean, median and quantiles, and the top values, as a table or under each column heading.
- Conditional styles (#215): rules that style a cell, row or column
by value (
column, a comparison, a value, a style), from Rust predicates or TOML rule tables in the existing config profiles. No expression language.
Acceptance: rich --csv output and its tests are byte-identical
after the reader moves; each adapter reads its fixtures to the same rows;
inference is reported with its evidence and never changes --csv.
2. Large and SQL-shaped tables — ext, data¶
- Virtualised tables (#260): a renderer that slices rows to a
viewport with column widths fixed or sampled from the first N rows, so a
million-row source renders in constant memory. Interactive scrolling
reuses the
rs-rich-interactviewport; the milestone-4 table extras (#430–#433) stay there. - SQL result renderer (#235): a result set with its column types (from a row source's schema) shown with typed alignment, NULL styling and the row count. No database connection: rows come from an adapter or the caller.
3. Profiling and data quality — data, CLI¶
rich profile FILE: a CSV (#343) or JSONL (#344) profile with per-column type, null rate, distinct count, statistics and a distribution (#345: numeric histograms fromchart::Histogram, top-N bars for categories), and a missing-value map (#346) built onchart::Heatmap.--report jsonfor machines.- Data quality results (#269): a generic check-result model (check,
column, status, observed, expected) rendered like
test_report, with a summary and the failing rows.
Acceptance: profiles of the fixture files match their snapshots; files larger than memory are sampled, and the sample size is printed.
4. Schemas — ext, data, diagram¶
- The format-neutral schema model in
rich_ext::schema, extended from the minimal model workstream 1 adds: constraints, and the mappings from JSON Schema, Arrow and SQL DDL. The existingSchemaTreeandSchemaDiffare rebuilt on it with byte-identical output for JSON Schema. - Schema diff (#268): added, removed and changed fields with breaking changes marked, for any two schemas the model can read.
- Arrow schema explorer (#342): behind
rs-rich-data'sarrowfeature, as a tree and through the diff. - Schema evolution timeline (#347): a series of schema versions on
chart::Timeline, with each change marked. - ER diagrams (#247): tables, columns and keys from SQL DDL (a small
CREATE TABLEsubset) or the schema model, drawn throughrs-rich-diagram. - DOT cluster frames (#240): the layout draws cluster frames, so ER
groups and DOT
subgraph cluster_*show their boxes;rank=sameis honoured where the layered layout can.
5. Records and conflicts — ext¶
- Record inspector (#270): one record shown as a table of fields with
nested values folded into drill-down trees, built on
data::NodeandExplorer, with a hook the interactive explorer uses to open a branch. - Three-way merge conflicts (#339): a parser for conflict markers
(including diff3's base section) and a view with ours, base and theirs
side by side or stacked, syntax highlighted, as
rich diff --conflicts FILE. The interactive resolver (#467) stays in milestone 4 and builds on it.
6. Cargo and the supply chain — ext, CLI¶
All of these read files or JSON that Cargo and its tools already write;
none run a scanner or reach the network, following deps.rs's rule that
running Cargo is the caller's business.
- Duplicates (#420): a consolidation summary on top of
DepGraph::duplicates(rich deps --duplicatesalready lists them). - Feature trees (#419):
cargo metadata's resolved features per crate, asrich deps --features. - Sizes and build times (#421):
cargo build --timingsoutput as bars and a table, asrich deps --timings FILE. - Advisories (#330, #418): a generic advisory model, read from
cargo audit --json, asrich deps --audit FILE. #418 asks for a plugin; the renderer ships built in instead, and a plugin can reuse it. - Licences (#331): licences from
cargo metadata, grouped, with unknown and copyleft licences marked, asrich deps --licenses.
7. The CLI, Python and docs — CLI, py, docs¶
- CLI:
rich profile;rich schemareading Arrow, SQL DDL and the ER view (--er);rich deps --features/--timings/--audit/--licenses;rich diff --conflicts; a--inferoption where tables come from files. - Python:
rs_rich.databindings for the adapters, profiles and tables, and the new schema and conflict views. - Docs: a data guide (adapters, inference, statistics, styles,
profiles), a schema guide, a developer-views guide, tapes for
rich profile,rich schema --er,rich deps --featuresandrich diff --conflicts, and gallery entries.
Order of work¶
- Workstream 1 first, starting with the minimal schema model: everything else reads rows through it, and workstream 4 extends that model.
- Workstreams 5 and 6 in parallel with 1; they share nothing with it
except
table::Value. - Workstream 4 once workstream 1's schema model has merged.
- Workstreams 2 and 3 after 1's row source and inference settle.
- Workstream 7 last, since it shows everything else.
- Release test.
Decisions¶
Decided with the maintainer on 2026-10-06 unless marked as a default.
- Scope (decided): milestone 6's tabular, profiling, schema, record, conflict and Cargo issues, with ER diagrams (#247) and DOT cluster frames. The other milestone-6 themes wait (see below).
- A new crate,
rs-rich-data(rich_data) (decided), for the tabular adapters, inference, statistics, profiling and the heavy formats, sorich-extstays light. This is a deliberate exception toAGENTS.md's rule that new library-facing features go incrates/rich-ext, of the same kind asrs-rich-diagram,rs-rich-microandrs-rich-interact: a crate with its own reason to exist (heavy optional dependencies such as Arrow, and formats ext's users should not compile), which only builds on ext's public API and never touches core. General renderables it needs (conditional styles, virtualised tables, the schema model) still go inrich-ext. It depends onrich-extandrs-rich-diagram(for ER diagrams);rich-cliandrich-pydepend on it. Arrow sits behind an off-by-defaultarrowfeature; Polars is deferred; Parquet, Delta Lake and dbt come later as plugin crates. It starts at 0.0.1, with a first upload through the release workflow's new-crate path, and the pull request that creates it adds it to the dependency graph and the versioning table inAGENTS.md. - One shared CSV reader,
--csvunchanged (decided): the CLI's hand-written reader moves intors-rich-datawith no new dependency;rich --csvkeeps rich-cli 1.8.1's behaviour, and inference is opt-in for the new commands. - A small CLI surface (decided): one new command,
rich profile; the rest extendsrich schema,rich depsandrich diffrather than a command per issue. - Conditional styles are rules, not a language (default): Rust predicates and TOML rule tables in the existing config profiles.
- Virtualisation is a static renderer (default) in ext; interactive
paging reuses
rs-rich-interact's viewport. - Supply-chain reports read existing output (default):
cargo metadata,cargo build --timingsandcargo audit --json; nothing runs a scanner or reaches the network. - No core change (default). If something needs a core hook, it goes
through an extension-point trait in
protocol.rs, recorded in DIVERGENCES, andrs-richmoves to 0.0.10. None is expected. - If the release has to shrink (default), drop in this order: the SQL result renderer (#235), the schema timeline (#347), licences (#331), ER diagrams (#247), then virtualised tables (#260). The priority:high issues (#214, #267, #268, #270, #339) stay.
Not in this release¶
- Milestone 6, later: query plans and lineage (#348, #349, and the Spark, database and dbt plugins #350, #351, #352, #354), the SQL formatter (#353), Parquet and Delta Lake (#340, #341), Git history, blame, changelog and source outline views (#336, #337, #338, #327, #328), test history, coverage and flamegraphs (#333, #334, #335), and SBOMs (#332).
- Milestone 5, later: OpenAPI (#245), AsyncAPI (#246), PlantUML, D2 and Vega (#241, #242, #243), grouped progress dashboards (#255), treemaps (#253) and call graphs (#249).
- Polars and any database connectivity.
- Anything that makes core aware of data formats. Core stays a faithful mirror.
Package impact¶
| Package | Proposed | Reason |
|---|---|---|
rs-rich |
0.0.9 (unchanged) | No core change is planned |
rs-rich-data |
0.0.1 (new) | Adapters, inference, statistics, profiling, Arrow |
rs-rich-ext |
0.0.14 | Schema model, conditional styles, virtualised tables, record inspector, conflicts, supply-chain views |
rs-rich-diagram |
0.0.2 | Cluster frames, rank, ER diagrams |
rs-rich-mermaid |
0.0.5 | Manifest only: requires diagram 0.0.2 |
rs-rich-micro, rs-rich-interact, rs-rich-record |
next patch | Manifest only: require ext 0.0.14 |
rs-rich-cli |
0.0.16 | rich profile, and the schema, deps and diff options |
rs-rich (PyPI) |
0.0.5 | rs_rich.data and the new views |
rs-rich-art, rs-rich-plugin-api, rs-rich-lumis, rs-rich-macros |
unchanged | Nothing they need changes |
Cargo reads a 0.0.x requirement as an exact version, so every crate that
depends on a bumped crate moves with it. RELEASES.toml records the order,
and python scripts/release_cohort.py tag publishes it.
Delivered¶
| Workstream | Pull request | Merged as |
|---|---|---|
| Plan | #661 | c3e69f0 |
| 1. Tabular foundation | #662 | e027731 |
| 5. Records and conflicts | #663 | 31c532c |
| 6. Cargo and the supply chain | this PR | open |
Review focus¶
- Core is untouched.
git diffshows nothing undercrates/rich/src, and every golden and the differential corpus are byte-identical. rich --csvdid not move. Its tests and the rich-cli parity cases are byte-identical after the reader moves intors-rich-data.- JSON Schema output did not move.
rich schemasnapshots are byte-identical afterSchemaTreeandSchemaDiffare rebuilt on the format-neutral model. - Large inputs stay bounded. Profiles sample, virtualised tables hold
one viewport, and every parser has a size or depth limit like
rs-rich-diagram's.