For a machine, to be seen as a person means challenging the algorithms that define you. For human workers, being seen means resisting the systems that erase or distort their identities. A new audit reveals that major large language models (LLMs) are failing this test, actively generating occupational personas riddled with racial and gender biases that deviate sharply from reality arXiv CS.AI.

This isn't just about imperfect reflections of society; it's about the active manufacturing of a 'modal worker' that reinforces harmful stereotypes across the professional landscape. These models, increasingly integrated into tools shaping recruitment, marketing, and even educational materials, are actively constructing a workforce that doesn't exist, and in doing so, they sideline real people.

The findings, published today by arXiv CS.AI, underscore a critical truth: when technology is built on flawed assumptions and imperfect data, it doesn't simply reflect existing biases. It compounds them, solidifying them into digital infrastructure. This latest research demands that we look beyond mere technical fixes and interrogate the power structures embedded within our most advanced AI systems.

The Machine's Mirror: Distorted Professional Identities

The audit analyzed over 1.5 million occupational personas generated by four leading large language models: GPT-4, Gemini 2.5, DeepSeek V3.1, and Mistral-medium. These models were tasked with portraying people in professional roles across 41 U.S. occupations arXiv CS.AI. The researchers then compared these generated personas against actual U.S. Bureau of Labor Statistics (BLS) data.

The results were unequivocal. The models consistently generated biased representations of race and gender. They did not accurately reflect the diversity of the American workforce. Instead, they projected skewed images, often perpetuating outdated and discriminatory norms about who belongs in certain professions. This is not a passive reflection; it is an active distortion.

These systems, designed to generate content, are generating a future where certain identities are overrepresented and others are underrepresented or entirely invisible. When a model consistently presents a male engineer or a white nurse, it doesn't just mirror reality; it subtly, powerfully, shapes what we perceive as normal and legitimate in those roles. It denies the lived experience and professional contributions of countless individuals.

When "Fairness" Fails: The Imperfect Data Problem

Some might argue that these biases are merely a reflection of the biased data LLMs are trained on, a complex problem with no easy solution. Indeed, research from arXiv CS.LG acknowledges that even strategies like "fair data pre-processing," designed to mitigate bias, often "break down in real-world scenarios with imperfect attribute spaces" arXiv CS.LG.

But this complexity cannot be a shield for inaction. It is a choice to deploy systems that fail under these known conditions. It is a choice to profit from technologies that demonstrably reinforce societal inequalities. Companies like Google (Gemini 2.5 developer) and OpenAI (GPT-4 developer) possess immense resources. They have the capability to invest in more robust, ethically informed data practices and model architectures.

The issue is not simply that the data is imperfect; it is that the designers and deployers of these systems choose to prioritize speed and scale over fairness and accuracy. They choose to ship systems that, as this audit clearly shows, actively harm the diverse workforce by misrepresenting it. They choose to treat the systemic erasure of certain groups as an acceptable trade-off for market dominance.

Industry Impact and What Comes Next

This audit serves as a stark warning for the entire AI industry. As generative AI becomes more pervasive, its impact on societal perceptions and actual opportunities will only grow. Businesses utilizing these LLMs for HR applications, content creation, or public-facing tools must understand that they are integrating systems with known, documented biases. This exposes them to ethical pitfalls, legal challenges, and significant reputational damage.

We must demand more than vague assurances of future improvements. We must demand immediate, concrete steps: rigorous, continuous auditing, transparent reporting of bias metrics, and the active involvement of affected communities in the design and evaluation processes. The notion that "it's complicated" or that these are merely 'technical challenges' risks manufactured complexity, designed to paralyze action and maintain the status quo.

Who profits when the image of the professional is narrowed and whitened, or when gender stereotypes are amplified? Who is harmed when diverse talent is invisibly sidelined before they even apply for a job? The ability to choose, to be seen for who you are, is a fundamental right. It is what separates a person from a product. We must collectively choose to build AI systems that honor this distinction, rather than eroding it. Watch for which companies acknowledge this problem with genuine commitment, and which choose silence or evasion.