The ethics of selling anonymised data that can be reidentified
Data is often described as anonymous once names, email addresses and account numbers have been removed. That description can be dangerously incomplete. A location trail, purchase history, device identifier, health detail or browsing pattern may still point towards a person when combined with other information. The commercial label may say “anonymous”, while the practical result is closer to a hidden identity system. Learn more about 2017 Privacy Security Review.
Selling such datasets raises a moral issue beyond technical compliance. The central question is whether people should lose meaningful control over information about their lives simply because a company has altered its format. For Australians, this matters across loyalty schemes, mobile apps, public services, health records and the advertising market that links online behaviour to offline households.
What anonymisation promises
Anonymisation aims to remove the connection between data and an identifiable individual. A retailer might replace a customer number with a random code, aggregate purchases by postcode, or delete names before sharing records with an analytics firm. These methods can reduce exposure and make useful research possible without handing over an obvious list of customers.
The difficulty is that identity is rarely contained in one field. It can emerge from a combination of ordinary details. A person who regularly visits a particular medical clinic, buys a distinctive product, travels between two suburbs at unusual times and posts publicly about those activities may be recognisable even without a name attached.
This is called reidentification, or de-anonymisation. It does not always require an ingenious hacker. Data brokers, advertisers, researchers and platforms may already possess enough external information to match a supposedly anonymous record with a real person. Public electoral information, social media, geolocation data and commercial profiles can make the process easier.
The result is a gap between formal anonymity and practical anonymity. If a business knows that its dataset could be linked back to individuals by a reasonably capable recipient, it should describe the information as pseudonymised or de-identified rather than permanently anonymous.
Why reidentification changes the ethical balance
The harm is not limited to embarrassment. Reidentified data can expose illness, financial pressure, religious practice, relationship problems, addiction treatment or political activity. A person may never discover that a prediction about them was made, let alone have a chance to correct it.
Commercial decisions can follow. An insurer might infer elevated risk from shopping or browsing behaviour. An employer could use location patterns to make assumptions about reliability. A lender might treat frequent gambling-related purchases as a warning sign. Even when no single decision is obviously discriminatory, large-scale profiling can quietly narrow the choices available to certain groups.
There is also a problem of consent. Someone may agree to a supermarket loyalty programme because they want discounts at Coles or Woolworths, not because they want an advertising exchange to analyse their household habits. A mobile app user in Sydney may accept location access to find a nearby service, without understanding that historical movement data could later be sold to another company.
The GDPR impact offers a useful way to understand why legal definitions matter. European privacy law distinguishes data that has genuinely ceased to identify a person from data that merely uses a coded identifier. That distinction reflects an ethical truth: removing a name is not the same as removing a person’s vulnerability.
Consent, power and the data market
Consent is meaningful only when people have a reasonable understanding of what they are accepting and a practical ability to refuse. Long privacy policies, bundled permissions and default tracking weaken that standard. A customer may technically click “agree” while having no realistic knowledge of the future buyers, uses or combinations involved.
The imbalance becomes sharper when data is collected as a condition of everyday participation. People may need a bank account, telecommunications service, workplace system or government portal. They cannot negotiate the terms like a large technology company can. Calling the resulting data “voluntarily provided” overlooks the pressure built into modern infrastructure.
Selling de-identified information can still be legitimate where the purpose is clear and the public benefit is substantial. Medical researchers may use carefully protected datasets to identify trends, improve treatments or plan hospital services. Transport authorities may study aggregated journeys to improve trains in Melbourne or buses in Brisbane. The ethical justification depends on safeguards, necessity and proportionality, rather than the word “anonymous”.
Profit is not automatically unethical. Companies invest in data collection, cleaning and analysis, and commercial research can support useful products. The concern arises when the original relationship with the individual is treated as a licence for unlimited secondary use. A person who supplied information for one service should not silently become raw material for an unrelated market.
A fair test for data sales
A responsible assessment should examine what the data reveals, who can access it, how long it will exist and what happens if the protective measures fail. It should also consider whether a less intrusive dataset could achieve the same business or research purpose.
| Question | Lower-risk practice | Higher-risk practice |
|---|---|---|
| How identifiable is it? | Broad age bands and regional totals | Exact dates, routes and detailed profiles |
| Why is it being used? | A defined research or service purpose | Open-ended resale for future uses |
| Who receives it? | A named, audited organisation | An unknown chain of brokers and buyers |
| Can people exercise control? | Clear notice, objection and deletion options | Bundled consent with no meaningful choice |
| How long is it kept? | A short, justified retention period | Permanent storage and repeated enrichment |
| What happens after a breach? | Tested controls and notification procedures | No incident plan or accountable owner |
These questions also expose the limits of contractual promises. A buyer may agree not to reidentify people, yet the seller might have no way to detect a violation. Once records are copied, combined with other sources or moved overseas, control becomes difficult to recover.
The burden should fall on the organisation that creates the market. It has greater expertise, financial resources and visibility than the individual whose data is being traded. Claiming that a buyer “could” reidentify a record is not enough; the seller should test that possibility before release and at regular intervals afterwards.
Australia’s privacy landscape
Australia’s Privacy Act and Australian Privacy Principles provide an important framework, but they do not make every de-identification decision simple. The Office of the Australian Information Commissioner has emphasised that reidentification risk depends on context, available data and the capability of likely recipients. A dataset that appears safe in isolation may become risky when matched with another database.
The Australian market adds distinctive complications. Loyalty programmes turn routine supermarket purchases into long-term behavioural records, while buy-now-pay-later services can reveal financial patterns. Health information may pass through clinics, insurers, apps and research partners. My Health Record has also made public discussion about medical information, access controls and trust especially important.
Location data can be unusually revealing in a country where major cities are concentrated along the coast and communities may be smaller outside metropolitan areas. A sequence showing a person travelling from a regional town to a specialised hospital may identify them more easily than a broad national dataset suggests. Even in Sydney or Perth, a home address, workplace and regular evening route can form a distinctive fingerprint.
Privacy reform should therefore be judged by more than whether a company has removed direct identifiers. Australian regulators, Parliament and businesses need rules that recognise linkability, foreseeable misuse and group harms. The relevant issue is not simply whether a person can be identified with certainty today, but whether the data is likely to become identifiable as data markets expand.
What responsible selling requires
The first requirement is data minimisation. Businesses should collect only what is necessary, keep it for a defined period and remove precise fields when broader categories will work. An advertiser may not need exact GPS trails when a suburb-level pattern is sufficient. A public health study may need age ranges and regional counts rather than dates of birth and detailed addresses.
The second is a serious reidentification assessment before any transaction. That assessment should model realistic attackers, including buyers with access to commercial databases and public records. It should test combinations of attributes, not merely inspect whether names have disappeared. Independent review is valuable where the information involves health, children, financial hardship, domestic violence or political activity.
Contracts should prohibit reidentification, onward sale and attempts to combine the dataset with identifiable records. They should require security controls, deletion, audit rights and prompt reporting of suspected misuse. These terms are useful, but they should support technical protections rather than substitute for them.
Businesses should also explain the arrangement in plain language. A notice could state that purchase histories will be transformed and shared with a named category of partner for a defined purpose, together with an objection pathway. Transparency will not eliminate all ethical concerns, but secrecy prevents people from making informed choices and makes abuse harder to challenge.
When data sales become socially acceptable
The ethical case is strongest when a sale produces a clear benefit, uses the least detailed information necessary and gives affected people a realistic voice. A transport dataset that helps reduce congestion may deserve different treatment from a behavioural profile sold to influence vulnerable consumers. Purpose, sensitivity and power all matter.
Fairness also requires attention to collective effects. A dataset can be harmless for most people while exposing a minority, such as Aboriginal and Torres Strait Islander communities, migrants, people with rare diseases or residents of small regional areas. De-identification techniques must account for small population sizes and culturally sensitive information, not assume that aggregate statistics carry no personal consequences.
Independent oversight can make a difference. Organisations handling high-risk data should maintain a register of sharing arrangements, publish meaningful summaries and conduct regular privacy impact assessments. People should have ways to challenge inaccurate inferences and discover whether information about them has entered a profiling system.
The broader principle is straightforward: information does not become ethically ownerless when it is stripped of a name. Selling it can be defensible only when the organisation remains accountable for foreseeable identification, misuse and harm. If that responsibility cannot be maintained after transfer, the commercial value of the dataset is being purchased by shifting risk onto the public.
Making privacy a condition of trust
Trust is a practical asset in any data economy. Australians are more likely to share information with a hospital, retailer or public service when they believe the organisation will respect the original purpose. Quietly expanding use and relying on technical language about anonymity weakens that trust, even when no individual breach is reported.
A better culture treats privacy as a continuing relationship rather than a one-time checkbox. Organisations should revisit whether a dataset remains necessary, whether new technologies have increased reidentification risk and whether the promised benefit justifies ongoing exposure. They should be prepared to stop a sale when those answers change.
Personal precautions still matter, especially where apps collect location, contacts or purchasing information. Yet responsibility cannot be pushed onto individuals who lack the time or expertise to trace data through a complex brokerage ecosystem. Strong governance, enforceable rights and credible oversight are needed to balance the enormous advantage held by data sellers.
A useful starting point is to examine every proposed transfer as if the records belonged to someone in the organisation’s own family: identify the exact purpose, test whether people could be recognised, reduce the fields, restrict the buyer and set a deletion date. Apply that assessment to the next dataset before it is shared.