The Lack of Race-based Data Interferes with the Task of Making Health AI Work for All in Canada’s AI Strategy
Author(s):
Dr. Tesh W Dagne
Dr. Laleh Seyyed-Kalantari

Disclaimer: The French version of this text has been auto-translated and has not been approved by the author.
The Federal government’s announcement of Canada’s AI strategy marks a significant step forward in using AI for the benefit of Canadians. Canada’s “AI for All” national strategy, unveiled last month, expands on the first phase of the Pan-Canadian Artificial Intelligence Strategy of 2017, which Canada was the first in the world to launch. The new strategy identifies six pillars of focus and priority sectors for continued investment in the AI ecosystem. The first pillar, “Protecting Canadians and safeguarding our democracy,” emphasizes the need for safety against AI risks as a key precondition for creating opportunities for adoption through the most fundamental form of trust. This pillar recognizes that trust in AI will lead to a confident adoption if Canadians do not consider the technology dangerous or harmful to themselves, their families and their communities.
As one of the priority sectors identified in the strategy, Canada’s healthcare sector has the potential to stimulate innovation in AI-enabled health technologies due to the wealth of clinical data in its universal healthcare system. However, the health data landscape in Canada faces significant hurdles that limit its utility for health AI technologies. While the government attempts to address well-known challenges, such as Canada’s fragmented health data landscape, through legislative reforms and investment measures, one challenge related to the creation of trust in AI remains unaddressed: the apprehension of equity-deserving populations towards AI due to its particular capabilities and threats.
Many commentators and communities, including Black communities, have long called for the collection and responsible use of race-based data in the healthcare system as a key tool to dismantle structural racism. This call takes on heightened significance if the need to create trust in AI, as recognized in the AI strategy, is to be realized. There are already established risks of racial bias being built into clinical AI models, resulting in AI failures such as the inability to diagnose skin diseases in Black and brown patients or reinforcing harmful stereotypes in healthcare.
The absence of race-based data in datasets used for AI development creates new threats of perpetuating health disparities on two grounds: first, research indicates that some health AI technologies perform poorly for certain equity-deserving groups compared with the general public. The lack of disaggregated race-based data in Canadian data challenges AI fairness researchers’ ability to understand the reasons and provide solutions. Second, research also indicates that for reasons researchers do not understand, health AI tools trained on medical images such as X-rays and Magnetic Resonance Imaging (MRI) scans unexpectedly and accurately discern the patient’s self-reported race. Such AI tools, designed solely to help clinicians diagnose patients from images, can recognize a patient’s racial identity even when healthcare providers or human experts interpreting the images do not know it. According to the American Civil Liberties Union, such capability could be abused as a discriminatory tool in health care provision, or it may unintentionally direct worse care to equity-deserving communities without detection or intervention.
There have been widespread calls for race-based data to aid in identifying inequities in healthcare and implementing interventions to address them, particularly during the COVID-19 pandemic. These calls have faced some opposition due to concerns that data may be used against communities, given the historical misuse of data, data extraction, surveillance, and privacy risks. There are, however, emerging community-based solutions that address such concerns, as well as practices that use race-based data to improve health outcomes for communities.
In the era of AI, Canada’s reluctance to collect race-based data and make it accessible to researchers overlooks the specific risks AI poses and may affect communities’ attitudes toward AI tools in health care. The collection of race-based data is common in other jurisdictions, such as the United Kingdom and some States in the United States. If the Federal government’s promise in its AI strategy to unlock the power of Canada’s health data as a driver of AI innovations is to bear fruit, the availability of race-based data is key in understanding the risks of AI and implementing interventions to address them. Measuring the biases of AI models in the healthcare system affecting equity-deserving groups is the first step toward mitigating them. However, such a task is currently difficult due to a lack of policy support for race-based data in Canada’s healthcare system. Removing this barrier supports AI safety assessment of equity-deserving groups, addressing potential biases and moving toward safer AI for all.

