Getting Started#
This guide helps you get up and running with ACRO for statistical disclosure control.
What is ACRO?#
ACRO (Automatic Checking of Research Outputs) is a Python package that provides statistical disclosure control for research outputs. It wraps common analysis functions, automatically checking for potential privacy disclosures before outputs leave a secure data environment.
From v1.0 onwards, ACRO’s checking logic is driven by a formal ontology see How ACRO’s Ontology-Driven Architecture Works below for an explanation of what that means in practice.
Key Concepts#
Statistical Disclosure Control (SDC)#
SDC is the process of protecting confidential information in statistical data releases. ACRO implements principles-based SDC that:
Identifies potentially disclosive outputs
Applies mitigation strategies when needed (suppression or rounding)
Maintains detailed, auditable records of every decision
Supports human checker workflows rather than blocking researchers
Disclosure Types#
ACRO checks for several types of disclosure:
Identity disclosure When individuals can be identified from a table cell
Attribute disclosure When sensitive attributes can be inferred about identified individuals
Inferential disclosure When statistical inference reveals information (e.g. dominance)
Linked-table disclosure When releasing two tables together reveals more than either alone
Safety Thresholds#
ACRO uses configurable thresholds to determine whether an output is safe:
Minimum cell count Default: 10 observations per cell
P-ratio threshold Default: 0.1 for dominance
NK-rule Default: n=2, k=90% for concentration
Degrees of freedom Default: ≥10 for regression models
These thresholds are set by the TRE administrator in a YAML configuration file and are loaded when your ACRO session starts. Researchers do not change them directly; see Configuration for details.
Basic Workflow#
A typical ACRO session follows four steps:
import acro
import pandas as pd
# Step 1: Initialise an ACRO session
# suppress=True means unsafe cells are removed automatically.
# Use suppress=False to see warnings without removing cells.
acro = acro.ACRO(suppress=True)
# Step 2: Load your data and run analysis as normal
df = pd.read_csv("my_data.csv")
result = acro.crosstab(df.region, df.income)
# Step 3: Review the output and add any exceptions if needed
acro.print_outputs()
acro.add_exception("output_0", "I need this output because...")
# Step 4: Finalise writes an audit report for the output checker
acro.finalise("safe_outputs")
The suppress=True option tells ACRO to apply suppression automatically.
If you prefer to see all results and decide yourself, use suppress=False.
Mitigation Strategies#
ACRO supports two mitigation strategies:
Suppression#
Unsafe cells are replaced with NaN. When margins are requested, they are
recomputed after suppression so they do not leak the suppressed values.
acro = acro.ACRO(suppress=True)
Rounding#
All cell values are rounded to the nearest multiple of a configurable base (default: 5). Marginal totals are recomputed from the rounded inner cells.
acro = acro.ACRO()
acro.enable_rounding(base=5)
How ACRO’s Ontology-Driven Architecture Works#
Prior to v1.0, the list of checks applied to each analysis was hard-coded. This made it difficult to:
Add support for new analysis types without touching multiple files.
Produce auditable statements of which checks ran and why.
Keep the code aligned with evolving SDC best practice.
The new architecture solves this by reading the check rules from a formal StatbarnsSDC ontology at build time.
The Four Lookup Tables#
When ACRO is built, ontology_handler.py reads the ontology and produces
four JSON files that are bundled with every release:
File |
Contents |
|---|---|
|
Maps analysis names (e.g. |
|
Maps each statbarn to the risks associated with it. |
|
Maps each risk to the checks that detect it and the mitigations that address it. |
|
Maps each check to the evidence it requires (e.g. a count table, residual degrees of freedom). |
Because these files are bundled with the package, ACRO works entirely offline inside a TRE no internet access is required.
What Happens at Runtime#
Session start
ACRO()creates anSDCChecksinstance which loads the four JSON files and the TRE’s risk appetite from the YAML config.Analysis call When you call, say,
acro.crosstab(...), ACRO:Creates a
TableModelDetailsobject holding all the parameters needed to reproduce the table.Looks up the appropriate analysis name (
"FrequencyTable","Mean", etc.) inanalyses.jsonto find its statbarn.Follows the chain: statbarn → risks → checks → evidence to determine exactly what data needs to be collected.
Collects all required evidence into an
SDCEvidenceobject.
Check and output Each required check runs on the collected evidence and returns a status (
pass,review, orfail) plus a plain-English summary. The results are combined and the chosen mitigation is applied.
Implementing support for a new type of analysis means:
identifying the type of analysis from the StatbarnsSDC ontology
creating a new ACRO function (typically as a method in
acro_regression.pyfor statsmodels-based models oracro_tables.pyfor pandas-based models)calling the package that creates the model inside your new function
adding three lines to:
specifying the name of the type of analysis
collect the evidence needed for disclosure risk assessment
process the evidence
see e.g. the regression functions in
acro_regression.py
Federated Mode#
ACRO also supports a federated mode for use with a trusted aggregator:
acro = acro.ACRO(federated=True)
In federated mode the evidence collection (step 2 above) still runs
locally inside the TRE, but the checks (step 3) are performed by a
remote trusted aggregator rather than locally. The evidence is packaged
and serialised by records.finalise_evidence(), which writes each
interim table to a CSV file and produces an evidence.json manifest.
This separation makes it possible to run SDC checks on aggregated evidence from multiple TREs without sharing individual-level data.
Installation Requirements#
Python 3.10 or higher
pandas, statsmodels, tabulate, PyYAML (installed automatically with
pip install acro)
Next Steps#
See Core Concepts for the detailed SDC methodology
Check Configuration for customisation options
Visit Examples for hands-on tutorials