When Public Data Stops Feeling Public
A dataset can be legally accessible and still be ethically sensitive. Electoral rolls, transport records, research repositories, property information, social media posts and government statistics often describe real people, even when names and email addresses have been removed. Once these records are combined, analysed or sold, individuals may become visible in ways they never expected.
The ethics of using public data sets that contain personal information therefore sits between privacy law, public interest and practical restraint. For Australians, the question is especially relevant as government open-data programs, commercial data brokers, loyalty schemes and artificial intelligence systems expand. A record being available to download does not automatically make every possible use fair.
Public Access Is Not the Same as Permission
People publish information in different settings for different reasons. A business may appear on an Australian Securities and Investments Commission register because transparency serves the public, while a person’s address may appear in a property record because the law has not caught up with modern risks. A photograph posted publicly on Instagram may be intended for friends, not for facial recognition training or a commercial profile.
This difference is often described as contextual integrity: information should flow in ways that fit the context in which it was shared. An open dataset can therefore create an ethical problem when its users ignore the original setting. A researcher studying rental affordability in Sydney may need aggregated suburb-level information. They do not automatically need a downloadable file containing precise addresses, household details and unusual personal circumstances.
The same principle applies when data is collected through scraping. Scraping a public website may be technically simple, yet the scale of collection can alter the nature of the information. A handful of public comments and a database of millions of comments are not ethically equivalent. Automation can turn scattered observations into a durable surveillance system.
This is why public-data decisions need more than a legal checklist. The questions should include whether people could reasonably anticipate the use, whether the purpose is proportionate, and whether the benefit could be achieved with less identifying detail. Privacy commentary at Twenty of Time examines this wider relationship between technology, power and personal autonomy.
De-Identification Has Limits
Removing names is a weak form of protection when other attributes remain. Birth year, postcode, occupation, dates of hospital visits and rare events can combine into a distinctive signature. A determined analyst may match that signature against news reports, electoral material, social media or another commercial database. This is known as re-identification, and it can happen without a dramatic security breach.
Australian organisations often treat de-identification as a technical exercise rather than an ongoing risk assessment. A dataset may be safe when released, then become easier to identify as new information appears elsewhere. The Australian Bureau of Statistics can publish useful population insights at a broad level, but a finely detailed release about a small community carries a different risk from national statistics. Small numbers can expose individuals through deduction.
Anonymisation should therefore be tested against realistic attackers, available auxiliary data and the consequences of exposure. Pseudonyms, hashed identifiers and random participant codes may help with internal management, but they are not automatically anonymous. If a party holds the key or can connect the code to another dataset, the information remains personal in practical terms.
Data minimisation is usually more reliable than relying on de-identification alone. Collect fewer fields, generalise locations, suppress rare categories, limit retention and restrict access to people who need it. Differential privacy and secure research environments can add technical safeguards, though they require expertise and careful calibration. A small loss of analytical precision may be a reasonable price for preventing personal harm.
Public Data Can Become Commercial Surveillance
The Australian data market makes the boundary between public information and private profiling particularly blurred. Coles and Woolworths loyalty programs can reveal purchasing patterns, while location data, app identifiers and advertising trackers help companies infer income, health interests or family circumstances. A public planning document may then be combined with consumer data to create a profile that the individual never knowingly supplied.
Vehicle systems provide a useful example. Modern infotainment platforms can gather location histories, paired-phone information, voice interactions, driving behaviour and device identifiers. A record produced inside a car may eventually be treated as a commercial asset rather than as a personal trace. The discussion of car data markets shows why “publicly available” and “fair to exploit” are very different standards.
Commercial use becomes more troubling when people cannot see the profile, correct errors or refuse secondary uses. An insurance company, landlord or employer could rely on an inferred characteristic without telling the person where it came from. A dataset may also reproduce social inequalities: suburbs with more monitoring generate richer records, while marginalised people face more scrutiny and fewer opportunities to challenge conclusions.
Ethical use requires purpose limitation and accountability across the whole chain. A university, government department or start-up should identify who will receive the output, what decisions may be influenced, and how affected people can contest mistakes. Selling access to a dataset does not transfer responsibility. Each recipient must assess whether its use is appropriate, secure and compatible with the original collection context.
Health and Community Data Need Special Care
Medical information carries an unusually high risk of embarrassment, discrimination and emotional harm. Even when a hospital dataset removes names, combinations of age, diagnosis, treatment date and regional location may identify a person with a rare condition. The problem becomes sharper in small Australian towns, where a handful of patients may be recognisable to local staff or neighbours.
Research can produce substantial public benefits, yet those benefits do not erase the need for governance. People who consent to treatment may not expect their records to be sold to an analytics company, used to train a general-purpose model or linked with prescription and consumer data. The limits of supposed privacy in anonymised medical data are important whenever institutions describe a database as risk-free.
Consent should be specific enough to be meaningful, while recognising that future research questions cannot always be predicted. Broad consent can be ethically defensible where there is independent oversight, clear withdrawal procedures and strict limits on commercial exploitation. Data custodians should publish plain-language explanations, conduct privacy impact assessments and involve patient representatives in governance rather than treating ethics approval as a one-time administrative hurdle.
Australian practice also needs to respect Indigenous data sovereignty. Information about Aboriginal and Torres Strait Islander peoples can affect families, communities and cultural authority, not merely individual participants. Community-controlled organisations and Indigenous governance bodies should have a genuine role in deciding how data is collected, interpreted, shared and destroyed. A technically accurate dataset can still cause harm if it strips knowledge from the community that generated it.
Responsible Use Requires Proportionality
A practical ethical test begins with the purpose. Is the project addressing a serious public need, improving transport, evaluating policy or merely making targeting more efficient? A public-health study examining heat exposure in Melbourne may justify carefully protected location data. A company wanting to infer which residents are likely to be vulnerable to aggressive marketing needs a much weaker justification.
The next test concerns necessity. Analysts should ask whether aggregated statistics, synthetic data or a smaller sample would answer the question. Precise GPS points may be unnecessary for a Brisbane traffic study; suburb-level counts could be enough. Similarly, publishing a full timestamp may add little value when a week or month would support the analysis. Limiting detail reduces the chance that people can be singled out.
Fairness must be assessed before release and after deployment. A model trained on public complaints may underrepresent people who lack reliable internet access. A dataset based on police interactions may reflect enforcement patterns rather than actual offending. Results should be tested across groups, documented with known limitations and reviewed when used in high-impact decisions such as housing, employment, insurance or access to services.
Good governance also includes security, deletion and redress. Access logs, tiered permissions, encryption and contractual restrictions can reduce misuse, but people need a way to report harm and correct records. The Australian Privacy Act and guidance from the Office of the Australian Information Commissioner provide important foundations, yet ethical conduct often demands more than minimum compliance. Organisations should be prepared to pause a project when the likely benefit is modest and the possible harm is difficult to reverse.
Public participation improves these decisions. Consultation should happen before data is released, in language people can understand, with attention to disability, language barriers and unequal digital access. A notice buried in a website policy is rarely enough. Trust grows when institutions explain what they collected, why they need it, who can access it and when the information will be deleted.
A dataset can serve the public without treating individuals as raw material. The fact that information sits online, in a government register or in a research archive establishes availability, not moral permission. The principle to remember is simple: use the least identifying data necessary, respect the context in which it was created, and keep the people behind the records at the centre of every decision.