Diagnose / Remove / Prove

Rosa diagnoses bias, removes it, and hands you the evidence.

Bias hides in data, directly and through proxies. Rosa strips it out without changing your data's statistics, and every run leaves an immutable, verifiable record. Everyone else measures and manages bias. Rosa removes it.

Free to try · Run it live in the browser

The most studied case of bias in data

ProPublica: COMPAS flagged Black defendants as future criminals at roughly twice the rate of White defendants.

68% of the false-positive-rate
gap by race, removed
the original race false-positive-rate gap
removed by Rosa
left

downstream recidivism model

bias dial 0.111 → 0.020 distribution preserved
Full raw COMPAS · median, range 64–83% · See the methodology

Measured, sourced, reproducible

The numbers do the persuading.

68% median reduction in a recidivism model's false-positive-rate gap by race (ProPublica's metric), on full raw COMPAS - range 64-83% Full raw COMPAS, downstream model
0.686 vs 0.598 downstream R² with Rosa plus a model, vs the model alone (0.546 with a parity constraint) - reproducible via the portal Test 1 one-click Test 1 credit-risk experiment
A bias score you can't game. Rosa measures residual bias with a held-out permutation test - on data with no real bias, it reads zero. So a low residual is a proven result, not a number we tuned. Verification rigor
bit-for-bit distribution preservation in training and inference output, to float64 precision Validation suite

R² (R-squared) measures how well a model's predictions fit reality; higher is better. Reproducible. See the methodology.

Why Rosa is different

The only tool in the conversation that changes the dataset and proves the change.

Open-source fairness toolkits

Assess and mitigate at the data scientist's desk. A toolbox of metrics and algorithms you assemble and operate yourself. No evidence artefact, no service.

Governance and GRC platforms

Measure, monitor, and document bias across the enterprise. Useful paperwork, but the dataset leaves the platform exactly as biased as it arrived.

Rosa

Removes the bias at source, preserves the data's statistics, and proves the change with an immutable, hash-verifiable record of every run.

The product in three beats

Diagnose. Remove. Prove.

01 / DIAGNOSE

Find the bias

Point Rosa at a dataset and one protected attribute. It finds where bias is encoded, directly or by proxy (a feature that stands in for the protected attribute, like a postcode for race), and scores it.

02 / REMOVE

Remove it at source

Rosa transforms the data so downstream models cannot recover the protected attribute, while preserving every column's statistics.

03 / PROVE

Keep the evidence

Every run emits an immutable Run Manifest and a PDF report. The audit evidence is generated by doing the work, not written up afterwards.

EU AI Act · Article 10

High-risk AI must be examined for dataset bias, with mitigation in place. Article 10 applies from 2 December 2027.

Article 10 of the European Union Artificial Intelligence Act requires the training, validation, and testing datasets behind high-risk AI systems to be examined for possible biases, with appropriate measures to detect, prevent, and mitigate them. Rosa produces exactly that examination, the mitigation, and the evidence, in one run.

EUR 35M / 7% maximum penalties under the Act for the most serious breaches; data-governance obligations for high-risk providers carry up to EUR 15M or 3% of worldwide turnover EU AI Act, Article 99 · how Rosa maps to Article 10
ARTICLE 10 · FROM 2 DECEMBER 2027

Product as proof

Do not take our word for it. Run it.

The Customer Portal includes Test 1: the Apple Card gender-bias scenario, debiased in your browser in one click. Submit the dataset, watch the job run, download the debiased CSV, the PDF report, and the Run Manifest (an immutable, hash-verifiable record of exactly what Rosa did to the data).

Run Manifest verified
job_id
550e8400-e29b-41d4-a716-446655440000
mode
remove_bias_training
job_status
complete
timestamp_submitted
2026-06-10T09:14:02Z
timestamp_completed
2026-06-10T09:31:47Z
row_count
4,645
bias_columns
["race"]
input_hash
sha256:9f1c…e7a2
schema_hash
sha256:4b08…21cd
config_hash
sha256:d3aa…90f4
container_digest
sha256:71be…0c55
bias (pre)
0.111
residual_bias
0.020
artifacts
remove_bias_report.pdf, compas_fair.csv
Immutable. Written once per job, kept indefinitely. Illustrative values; field set mirrors the live manifest schema.

What sets Rosa apart

  • No proxy labelling required. Name one protected attribute; Rosa finds the proxies itself.
  • Any model or stack. Rosa outputs a clean CSV. No rewrite of your training code.
  • Data residency you choose. UK region for the trial; a dedicated instance in your own region for a PoC.
  • Works for human decisions too. Cleans the data analysts use directly, not just model training data.
  • No fairness/accuracy trade-off. Fairness as a transformation of the data, not a constraint fighting your model.
  • Honest scope. Univariate: one protected attribute per run. Rosa says what it can and cannot detect.

Designed to support organisations implementing these frameworks at the data layer

EU AI Act ISO/IEC 42001 NIST AI RMF SOC 2 GDPR EN 304 223

Rosa is an evidence-producing instrument, not a certification. Standards and compliance

Remove the bias from your data. Keep the evidence.

Free to try. Free for regulators, with no end date. Processed in the UK.