top of page

Has Research Ethics kept up with Data?

  • Aug 25
  • 6 min read

I have been thinking a lot about data recently, and the research ethics world that we now find ourselves in. Last week, our committee was fortunate to have a presentation by one of our members Pascarn Dickinson on Geolocation Data, which prompted my internal turmoil.


Then I went down the rabbit hole. For example, a voice recording used to be relatively easy to understand. A researcher interviewed someone, recorded the conversation, transcribed it and then most likely deleted it.


A recent article by Caroline Carver at Three Black Cats, From Conversation to Biometric Data, explores a problem when using AI note-taking tools. Something that appears to be simply recording and transcribing a conversation may actually be collecting information capable of much more sophisticated processing.


The data hasn't necessarily changed. What can be done with it has.


It raises a broader question for those of us working in research ethics: Have our approaches to data ethics kept pace with what data can now do?


The traditional data questions are still important, but are they enough?

Research ethics committees have long considered the management of research data. In absence of a data management plan, one might ask:


·      Who will have access to the data?

·      Where will it be stored?

·      What format will it be stored? Is it identifiable?

·      How long will it be retained?

·      Can it be used for future research?

·      Has the participant consented to this, and for future use?


Aotearoa New Zealand's National Ethical Standards address many of these issues in relation to health data, including secondary use, data governance, security, re-identification and the use of artificial intelligence (see NEAC Chapters 12 and 13). But increasingly, research data creates ethical questions that don't fit neatly within a traditional privacy-and-security model.


Consider a dataset that has been carefully de-identified and stored in an 'appropriately' secure environment.


That might answer the question:


Can we keep this data safe?


It doesn't necessarily answer:


Should these data be combined, transformed, analysed or used for this purpose in the first place?


That distinction is increasingly important.


Data can become more revealing over time

Research ethics has historically placed considerable emphasis on whether information is identifiable. But identifiability isn't necessarily a fixed property of a dataset. Particularly in a country like Aotearoa; we are a very small population, and two degrees of separation is truly a thing.


The ability to link datasets, improvements in computational power, artificial intelligence and increasingly sophisticated analytical techniques can change what can be learned from information that was previously considered relatively innocuous.


  • A voice can become a biometric identifier.

  • Images can support facial recognition.

  • Location information can reveal patterns of behaviour and association.

  • Administrative datasets can become considerably more revealing when linked with other information.

  • Large collections of free text can now be analysed at a scale and in ways that would have been impractical only a few years ago.


The ethical significance of data can therefore change even when the underlying information does not. This creates a rather difficult problem for research ethics review. We are making decisions about data today in an environment where its capabilities tomorrow may be very different.


Privacy compliance and data ethics aren't the same thing

Aotearoa has significant protections governing personal information. The Privacy Act 2020 establishes requirements around the collection, use, disclosure and protection of personal information.


More recently, the Biometric Processing Privacy Code has introduced specific requirements governing biometric processing, including questions of necessity, proportionality and safeguards.


These are important developments. But legal compliance and ethical acceptability aren't necessarily the same question. Privacy regulation might tell us whether particular information can lawfully be collected or used. Data ethics asks additional questions.


  • Is collecting the information necessary?

  • Is the proposed use proportionate?

  • What new information might be inferred?

  • Could combining datasets create risks that weren't present in either dataset individually?

  • Who benefits from the use?

  • Who carries the risk?

  • Are there impacts on groups or communities even where individual privacy is protected?


And importantly:


Would the people represented in the data reasonably expect it to be used in this way?


A dataset can be secure, lawfully held and technically de-identified while still raising legitimate ethical concerns about how it is being used.


Consent doesn't solve everything either

Consent remains an important part of ethical research. But contemporary data use can stretch what meaningful consent can reasonably accomplish.


Participants may agree that their information can be retained for future research. But how clearly can we explain future analytical capabilities that do not yet exist?


Likewise, someone whose information enters an administrative dataset may never have anticipated that it could later be linked with several other datasets to answer an entirely different research question.


Or what about me, with my geolocation information on my phone being bought and sold without my knowledge to outline my terrible spending habits and my caffeine addiction?


The answer cannot simply be to write increasingly broad consent forms. Good data ethics also requires governance, accountability and ongoing judgement about whether new uses remain consistent with the rights and reasonable expectations of the people and communities represented in the data.


Individual privacy isn't the whole picture

There is another limitation to focusing primarily on identifiability and individual privacy. Data can affect groups and communities.


Te Mana Raraunga has long argued that Māori data should not be understood simply as information relating to individual Māori. Māori Data Sovereignty recognises inherent Māori rights and interests in the collection, ownership, governance and application of Māori data.


That changes the ethical question.


Removing someone's name from a dataset does not necessarily remove its implications for their whānau, hapū, iwi or other communities. Nor does individual consent necessarily resolve questions about collective interests, governance, benefit and potential harm.


Contemporary data ethics therefore requires us to think beyond the individual participant and consider the relationships, communities and populations represented by data.


Perhaps we need to ask different questions

None of this means the traditional questions asked by ethics committees are wrong. We still need to know where data will be stored, who can access it, how it will be protected and whether participants have provided appropriate consent.


But perhaps contemporary ethics review needs another layer. When considering research involving data, we might also ask:


  • What can these data reveal now, and what might they reasonably reveal in the future?

  • What changes when these data are linked with other information?

  • What information might be inferred rather than directly collected?

  • Is the proposed use consistent with the context in which the data were originally provided?

  • Who has decision-making authority over future uses?

  • Who benefits from the research, and who carries its risks?

  • Could individuals, groups or communities experience harm even if nobody is personally identified?

  • Are there Māori rights and interests in the data that require appropriate governance rather than simply consultation?


These aren't solely questions about data security. They are questions about purpose, power, relationships, governance and responsibility.


Keeping pace

Technology will continue to change what researchers can learn from data. Our ethical frameworks cannot predict every future analytical technique, nor should we try to create rules for every possible technology.


But they can ensure that our thinking moves beyond a model in which responsible data use is primarily demonstrated through consent, de-identification and secure storage.

Those safeguards remain essential. They are simply no longer the whole conversation.


The challenge for research ethics is therefore not only to keep research data safe. It is to keep pace with what that data can become.


Additional reading

Caroline Carver – From Conversation to Biometric Data, Three Black Cats

The article that prompted this discussion, examining AI note-taking tools, voice data and the implications of New Zealand's new biometric privacy framework.


Pascarn Dickinson - The use and misuse of geolocation mobility data

A very important literature review: https://www.covid19modelling.ac.nz/geolocation-data/


Office of the Privacy Commissioner – Biometric Processing Privacy Code 2025

New Zealand's specific privacy framework for biometric processing, including requirements relating to necessity, proportionality, transparency and safeguards.


National Ethics Advisory Committee – National Ethical Standards, Chapters 12 and 13

The current national ethical guidance covering health data, secondary use, data governance, artificial intelligence, re-identification and emerging technologies.


Te Mana Raraunga – Principles of Māori Data Sovereignty

An important framework for understanding Māori rights and interests in the collection, ownership and application of Māori data, and why responsible data governance extends beyond individual privacy and consent.

 
 
 

Comments


bottom of page