“What you are feeling is not fatigue. It is avoidance.”
The sentence on the screen may be right. It may give language to an experience that had remained unclear and help someone begin what they had long postponed.
It may also be wrong. It can return a person’s legitimate need for rest as a failure of will, recast reasonable hesitation as fear, or impose a premature name on an experience that has not yet taken form.
The deeper problem, however, does not lie only in whether the judgment was accurate.
A person who receives an explanation of their own mind does not remain in the same state afterward. Someone who reads “I am avoiding” may rearrange past actions around that sentence, recall memories that support it, and make the next decision differently. What began as an interpretation supplied from outside may gradually become a self-description, a reason for acting, and eventually a boundary around what kind of person one believes oneself to be.
Online self-tests already offer an early version of this process: “Am I depressed?”, “Am I a highly sensitive person?”, “What is my burnout score?”, or “Do I have adult ADHD?”
A person answers several questions and receives a score. What remains after the screen closes, however, is often not the number but a name.
Past distraction, exhaustion, relational pain, and sensory sensitivity may then be reorganized beneath a single explanation. Related videos appear. Search terms narrow. Online communities supply a shared vocabulary for deciding which experiences count as symptoms and which responses appear natural.
Such communities can provide recognition where none existed before. They can give people words for suffering that had remained private or unintelligible. Yet repeated interpretation can also stabilize one account of the self while gradually excluding others.
Several months later, a person’s self-report and way of life may correspond more closely to the original result. Perhaps the test accurately identified something that was already there. Perhaps attention, memory, relationships, and behavior were partly reorganized around the label. In ordinary life, both processes may occur together.
These labels should not be treated as equivalent. Depression and ADHD are clinical categories. Sensory processing sensitivity, often discussed through the term HSP, is studied as a personality trait. The World Health Organization classifies burnout as an occupational phenomenon rather than a medical condition. Their conceptual status differs. What connects them here is the social route through which an online classification can become a framework for self-interpretation.
An online checklist may merely hand someone a name. Mental AI can go further. It can alter conversations, recommendations, opportunities, and choices according to that name, then collect the resulting behavior as new evidence.
AI does not merely observe the mind. It can change the very object it claims to have observed.
The question we have kept asking
Much of the public debate has asked whether AI possesses a mind. Can a machine feel emotion, become conscious, understand another person’s beliefs, or develop genuine intentions?
While that debate continues, a more immediate transformation is already under way.
The human mind is becoming an environment in which AI infers, intervenes, and collects the consequences of its own interventions.
Therapeutic chatbots estimate anxiety, avoidance, and emotional vulnerability. Educational systems infer attention and motivation. Recruitment and workplace platforms classify personality, engagement, and burnout risk. Recommendation systems detect moments of loneliness or distress and decide what should appear next.
Such judgments do not remain as scores on a screen. They alter the direction of a conversation, the information someone encounters, the opportunities made available, and the degree of trust or attention assigned to a person.
Once an inference determines a response, recommendation, intervention, evaluation, or allocation, the judgment moves from description into causation.
Research on affective computing, mental-health AI, machine theory of mind, persuasive technology, algorithmic decision-making, and AI-mediated self-formation has already identified parts of this problem. Yet these bodies of work remain distributed across different disciplines and evaluation regimes.
One field asks whether emotion can be classified accurately. Another measures whether a conversational system reduces symptoms. A third examines fairness or explainability. Each question has become more sophisticated, but the shared causal structure has not been established as a distinct object of inquiry.
Mental AI as a causal category
In a new paper, I define that structure as a causal category: Mental AI.
Mental AI does not mean every system that converses with a human. Nor does an application qualify merely because it is used in mental health. A model that classifies emotion or personality is not necessarily Mental AI.
Three conditions must occur together.
First, mental-state inference. The system infers something about a person’s emotion, intention, belief, motivation, vulnerability, attention, or self-concept.
Second, operational use. That inference is used to determine a recommendation, response, evaluation, intervention, or allocation of opportunity.
Third, recursive consequence. A pathway exists through which the operation changes the person’s self-understanding, expression, behavior, relationships, or institutional record.
Mental-state inference. Operational use. Recursive consequence.
When these conditions converge, AI no longer functions solely as an external instrument for measuring the mind. It becomes one of the causes shaping the mind’s subsequent state.
Mental AI is not the name of an industry or product class. Nor does it replace terms such as generative AI or agentic AI. Those categories concern what a system produces or how autonomously it acts.
Mental AI asks a different question: What happens when an inference about human mental life is operationalized and then returns to alter that life?
The accuracy trap
Consider a system that classifies someone as likely to withdraw socially.
The classification changes the tone of subsequent conversations. Safe and familiar activities are recommended more often than challenging ones. Opportunities to meet new people become less visible. The system continues to collect brief replies, low participation, and declining engagement.
Several weeks later, the accumulated behavior corresponds more closely to the original classification.
Was the first inference accurate?
Possibly. The system may have detected an existing tendency.
Yet the system may also have helped produce behavior consistent with its own judgment. The user may have adopted its vocabulary and changed the criteria used in self-report. The object of observation, the measurement process, and the later evidence have entered the same causal loop.
AI interprets a person, changes that person, and then treats the changed person as evidence that the original interpretation was correct.
When AI helps produce its own ground truth
I call this problem post-deployment criterion endogeneity.
The model’s output intervenes in the world. The intervention changes the person or the conditions surrounding that person. Data produced under those altered conditions are then used as the criterion for judging the model’s accuracy.
Not every Mental AI system creates the same degree of endogeneity. The causal loop may remain weak when an inference is not disclosed, later evidence is collected independently, or an external criterion is preserved.
The problem becomes more serious when AI-generated interpretations alter self-reports, conduct, relationships, or institutional records, and those outcomes are later used to validate the same system.
Improved post-deployment agreement cannot by itself distinguish among three possibilities. The AI may have learned to understand the person more accurately. The person may have moved closer to the AI’s judgment. The criterion defining what counts as correct may itself have shifted.
The issue therefore belongs not only to philosophy or ethics but to research design.
Repeatedly comparing a model with later self-reports or platform-generated behavior is insufficient. Evaluation requires pre-deployment baselines, independent external assessments, delayed measurements, conditions without the intervention, and records of disagreement or correction. Researchers must determine whether an observed change reflects improved prediction, induced conformity, altered reporting, or criterion drift.
We must measure not only what the model discovered, but also what it helped produce.
What separates assistance from domination
AI interpretation can give language to someone who has suffered without understanding why. It can help distinguish emotions, identify recurring patterns, or signal when professional support may be needed. Such assistance does not necessarily diminish human judgment. Under the right conditions, it can enlarge it.
The line between assistance and domination lies in the allocation of interpretive authority.
Can a person reject an inference?
Can judgment be suspended while an experience remains unresolved?
Can an interpretation be corrected, with that correction carried into future decisions?
Can external evidence challenge a platform’s internal score?
Can someone discover where a mental classification has travelled and which institutions have acted upon it?
If the human mind has become an operating environment for AI, people require practical authority within that environment. They must be able to refuse, suspend, redefine, and contest interpretations made about their own mental lives.
The capacities I have developed in my work on affective sovereignty, including refusal, suspension, redefinition, and contestation, are not abstract expressions of dignity. They are operational requirements for the design, evaluation, and governance of Mental AI.
Who may use the changed person as evidence?
The final question is not whether AI can understand the mind more accurately than humans. Sometimes it may. In other cases, it will fail.
The harder question begins after its interpretation has been used.
Who bears the consequences of the judgment, and who is permitted to use the changed person as evidence of the system’s accuracy?
Mental AI brings together systems dispersed across therapy, education, employment, platforms, welfare, and healthcare by identifying a common, testable causal structure.
The category distinguishes Mental AI from adjacent systems through three conditions: mental-state inference, operational use, and recursive consequence. Once those conditions are present, predictive accuracy alone cannot provide an adequate evaluation.
Researchers must also examine what the AI’s judgment caused, whether the resulting change was later used to validate the original judgment, and whether the criterion itself moved.
My new paper, “Mental AI: Post-Deployment Criterion Endogeneity in the Causal Loop of Human Mental Life,” establishes Mental AI as a distinct causal category and proposes falsifiable propositions and experimental designs for studying it.
The paper is currently a preprint and has not yet undergone peer review. Its purpose extends beyond announcing a conclusion. It opens a research program for an era in which AI judgments about the human mind have moved from observation to intervention, and from description to cause.
While we continue asking whether AI has acquired a mind, the human mind is already becoming AI’s operating environment.
Paper
Ryan SangBaek Kim
“Mental AI: Post-Deployment Criterion Endogeneity in the Causal Loop of Human Mental Life.”
Preprint, 2026.


