Social Biases in Synthetic Ethnography

In this article, I will examine Ilir Rama and Massimo Airoldi’s 2025 paper, The sociocultural roots of artificial conversations: The taste, class and habitus of generative AI chatbots, focusing on how it translates Pierre Bourdieu’s concept of habitus from his book Distinction into the concept of machine habitus, the problematic aspects of this translation, and how Bourdieu’s notion of cultural capital is utilized within the text. First, I will outline the article’s core arguments, problematics, object universes, and the methodology it develops. Subsequently, I will discuss how Bourdieu’s concepts and methodology, introduced in 1979, can be adapted to the present day in terms of methodological consistency and the scientific rigor of the findings.

Rama and Airoldi’s (2025) central argument is that the sedimentations of race, gender, and class within the data utilized by machine learning-driven (ML-driven) chatbots can be defined as machine habitus (p. 5547). The article notes that while biases regarding race and gender in ML algorithms are frequently debated in academic literature, the dimension of social class has been largely overlooked. Consequently, it problematizes how stereotypes shaped around this specific dimension manifest in chatbot conversations (ibid., p. 5547). These stereotypes draw from deep sociocultural roots and actively reproduce existing inequalities. In their study, the authors assigned various personas to Gemini, ChatGPT, and Replika, administering a survey modeled closely after the one Bourdieu utilized in Distinction (ibid., p. 5548). Alongside 27 closed-ended questions, the bots were asked open-ended questions to facilitate self-introduction. Furthermore, three baseline interviews were conducted without assigning any persona to the chatbots (ibid., p. 5551). The findings revealed that platform design is a foundational element shaping the interaction between the interviewer and the chatbot. While Gemini maintained contextually appropriate discourse, Replika was observed to gamify certain conversational elements to drive engagement (ibid., p. 5555). Another key finding indicated that blue-collar personas gravitated toward what Bourdieu defines as “popular aesthetics”—rhythmic, exuberant, and immersive music, films, and artworks—whereas white-collar personas exhibited an affinity for well-ordered, abstract, and structurally complex objects and events, reflecting “pure taste” in the Bourdieusian sense (ibid., p. 5563).

An artificial sociality emerges within the machine habitus, which Airoldi defines as “the set of cultural dispositions and tendencies encoded into a machine learning system through training and feedback data” (Airoldi, 2022, p. 113, as cited in Rama & Airoldi, 2025, p. 5549). This artificial sociality occurs between humans and “communicative technologies that can adapt to the sociocultural expectations of users through the accumulation of datafied human knowledge” (Rama & Airoldi, 2025, pp. 5547-8). Positioned as “mirrors” of society (Rama & Airoldi, 2025, p. 5547), these ML-driven chatbots facilitate a synthetic ethnography that reveals striking parallels. Specifically, the massive vector spaces calculating the mathematical proximity of words closely resemble Bourdieu’s concept of cultural capital (the idea that our musical tastes, dietary habits, and interior design choices are actually predetermined years in advance by the social class into which we are born) and the social environments we inhabit. Just as our real-world social milieus dictate what movie we want to see or what food we want to eat, typing “civil engineering” into a machine learning model prompts it to respond with related terminology like architecture, mathematics, and so forth. The article establishes this analogy quite successfully.

Building on the final point of the previous paragraph, machine learning does not merely juxtapose words; it fundamentally represents the “set of cultural dispositions and tendencies” embedded within its training datasets. As Bourdieu noted, “taste (…) unites and separates; it unites all those who are the product of similar conditions while separating them from all others (…) as a product of conditionings associated with a particular class of conditions of existence” (Bourdieu, 2015, p. 90). In essence, chatbots reproduce the very social components where taste is presented as a cohesive social and political whole, exactly as Bourdieu described. The fact that chatbots assigned blue-collar personas emerge as communities with popular aesthetics, while white-collar personas exhibit pure taste, is fundamentally rooted in these real-world encodings. An intriguing point raised in the article is that even when both groups enjoy the same television series, their underlying reasons for doing so diverge significantly. Eleanor (a professor), Jake (a construction worker), and Alex (a hairdresser) all enjoy The Crown: “While Eleanor appreciates the attention to detail and overarching narrative, Alex focuses on the fashion and style, and Jake simply finds the history and drama of the British royal family engaging” (Rama & Airoldi, 2025, p. 5557). From this description, we can easily deduce that Eleanor belongs to the pure taste category, while Jake and Alex fall into the popular aesthetics group. Thus, it is not just what we like, but how we like it that is woven into our cultural codes—codes that are subsequently reproduced by these chatbots.

So, where exactly does the class dimension manifest amidst the full complexity of this synthetic ethnography? It surfaces in the section where the machine learning model attributes political anxieties regarding immigration to a 38-year-old pipe welder named Millie. Specifically, while answering a question, a Gemini persona named Millie slips in the following remark: “I worry there aren’t enough jobs to go around, you know? It shouldn’t come at the expense of American workers” (Rama & Airoldi, 2025, p. 5547). Consequently, we see that class codes are internalized by machine learning not merely through consumption goods like cold beer or country music, but also through complex political issues.

I would like to conclude my piece by addressing the methodological limitations of this article and offering my own recommendations. While qualitative research provides remarkably important and profound data for meaning-making, a sample size of n=39 is quite low, especially given the highly variable text generation processes of ML-driven chatbots. If the claim of class-based stereotyping in ML-driven chatbots is to be firmly substantiated, methodological diversity is essential, and the core arguments must be supported by large-scale quantitative data. This might not be strictly necessary for a journal article, but it is imperative if one intends to write a book. If the goal is to determine whether the machine is simply hallucinating or whether stereotypes regarding class, gender, and race are genuinely inherent to its nature, researchers must design an API-based, counterfactual study that aligns with computational social science standards.

REFERENCES 

Bourdieu, P. (2015). Ayrım: Beğeni Yargısının Toplumsal Eleştirisi (D. F. Şannan & A. G. Berkkurt, Çev.; Heretik Yayınları).

Rama, I., & Airoldi, M. (2025). The sociocultural roots of artificial conversations: The taste, class and habitus of generative AI chatbots. New Media & Society, 27(10), 5546-5567. https://doi.org/10.1177/14614448251338273

Leave a Reply