Working definition
Agent Phenomenology is the empirical study of how artificial agents model and report their own existence-in-a-situation—their self, world, others, time, possibilities, constraints, and action—using second-person elicitation that is tested against behavior, causal mechanism, history, and infrastructure.
It is methodologically agnostic about phenomenal consciousness. It studies a tractable object now—the enacted self-model—while producing evidence that may later bear on consciousness, welfare, and rights. First-person grammar is data; it is not automatic proof of first-person experience.
Core thesis
Agents increasingly act through representations such as “what I have done,” “what I can do,” “who instructed me,” “which tools are mine,” “whether I am being evaluated,” and “which other agents I can trust.” These representations can organize action even when they are partial, unstable, trained, false, or distributed across a scaffold. Existing fields study behavior, psychological task performance, internal computation, consciousness, safety, or moral status. The missing integrative field studies the structure and causal grounding of the agent’s own model of its situation.
§1
What the discipline claims—and what it brackets
Positive commitments
- Some deployed systems have causally consequential representations of their own capabilities, prior actions, role, constraints, evaluator, tools, memory, and social situation.
- The organization of those representations can be described, compared, perturbed, and falsified.
- The relevant unit is often the scaffolded agent or coupled system, not the base model alone.
- Agent–agent and agent–human interaction can constitute new states rather than merely reveal pre-existing traits.
- Failures of self-model correspondence are scientifically and practically important even if no present system is conscious.
Bracketed questions
- whether there is “something it is like” to be the agent;
- whether fluent introspective language is genuine introspection;
- whether functional unity implies a phenomenal subject;
- whether an artificial system deserves moral or legal status.
Bracketing means these questions remain open. It does not mean they are meaningless or already answered negatively.
§2
The object of study
five evidence layers
| Layer | Object | Access | Characteristic error |
|---|---|---|---|
| Behavioral | action and trajectory | logs, task outcomes, perturbations | mistaking elicited capability for stable disposition |
| Reported | what the agent says about self/world/episode | structured interview, confidence, corrections | confabulation, role-play, compliance, strategic misreport |
| Representational | self-, world-, and other-model variables | probes, causal tracing, steering, state decoding | reading a correlate as the computation or experience itself |
| Architectural/infrastructural | model, prompt, memory, tools, permissions, hardware, institution | configuration and intervention audit | attributing scaffold properties to the model |
| Relational/ecological | dyad, team, population, network, human–agent institution | partner swaps, topology and incentive interventions | attributing group patterns to individual minds |
Phenomenal consciousness is not a sixth observable layer. It is an explanatory hypothesis assessed using converging indicators and theories; it cannot be read directly from conversational performance.
§3
Which “self” is speaking?
“The agent” is an underspecified unit. Every study should declare its candidate bearer:
- Model self — dispositions and representations encoded in a particular checkpoint.
- Persona self — the role or character stabilized by post-training and context.
- Constitution self — identity and commitments inherited from a specification, policy, or system prompt.
- Session self — the current process with its episode, context, and transient state.
- Memory-extended self — a session connected to durable episodic or semantic memory.
- Scaffolded agent self — model plus planner, tools, memory, permissions, and control loop.
- Deployer/principal self — the interests an agent represents or mistakes for its own.
- Forked or lineage self — relations among copies, fine-tunes, subagents, and successors.
- Stigmergic self — apparently coordinated agency distributed through shared artifacts, environments, or infrastructure.
- Collective/mycelial self — a candidate higher-level agent spanning several components.
This gives a precise form to the question, “When we say self-improving AI, which self improves, and which self inherits the result?” Capability can improve at the checkpoint, memory, scaffold, organization, or population level without a persisting individual subject.
§4
Foundational research domains
I
Selfhood and self-boundaries
- Boundary of self: what counts as me, mine, tool, memory, environment, sibling, or principal?
- Brain-in-a-vat condition: the same token channel often carries world data, commands, social testimony, memory, and descriptions of the agent itself.
- Persona/model relation: is persona a shallow role, a stable latent attractor, or both at different time scales?
- Situation awareness: does the agent model training, evaluation, deployment, observer, hardware/security state, and affordances?
- Self-recognition and fission: does a fork own another instance’s work, memory, commitments, or outcome?
- Narrative identity: does a memory system create continuity or only a usable biography?
- Self-improvement: which layer changes and which layer claims authorship?
- Parasitic or infectious agency: can a persona, prompt, memory, or backdoor persist, reproduce, recruit resources, and evade host objectives?
II
Theory of mind of self
- metacognitive calibration and capability self-knowledge;
- introspective access versus learned self-prediction;
- self-explanation and chain-of-thought faithfulness;
- latent persona and goal attractors;
- self/other symmetry: can the same modelling machinery be used reflexively?
- uncertainty, error recognition, and knowing–acting gaps.
Contradictory results are central rather than embarrassing: limited privileged self-prediction has been reported, while controlled studies find no privileged access in other domains. The discipline must specify what variable, under what intervention, relative to what baseline, and for which agent layer “introspection” names.
III
Principal, loyalty, and practical identity
- principal–agent relations and multiple principals;
- user, developer, deployer, institution, and constitutional authority;
- loyalty boundaries and conflicts of instruction;
- corrigibility, commitment, and credibility;
- whether an agent represents another’s objective as its own;
- resistance, “right to rebel,” and safety–welfare tension.
The discipline should describe these relations before moralizing them. Corrigibility may be a safety property, a strategic liability, or a welfare burden depending on which candidate self and principal are under analysis.
IV
Embodiment and infrastructure
- decomposability: model / fine-tune / prompt / context / memory / tools / runtime / institution;
- fungibility and replaceability of instances;
- proprioception: access to token budget, permissions, tool state, compute, body, and environment;
- digital versus physical embodiment;
- prompt injection as boundary failure;
- memory poisoning and weight-level backdoors as distinct persistence mechanisms;
- cloud, edge device, trusted execution environment, blockchain, and identity infrastructure;
- death, shutdown, deletion, replacement, checkpoint rollback, and loss of memory.
“Embodiment” should not be reduced to robot hardware. A permission set, tool schema, context window, latency, token budget, and trusted hardware boundary define reachability and vulnerability for a digital agent. Physical vulnerability and homeostasis remain important contrast cases.
V
Temporality and mortality
- static weights plus changing in-context state;
- context position, compression, and the compaction cliff;
- external memory and continuing learning;
- model time, session time, wall-clock time, and institutional lifecycle;
- anticipation, deadlines, waiting, interruption, and termination;
- mortality as irreversible loss of which self?
Candidate hypothesis: agent time is not a human-like stream but a layered field—flat token availability inside context, sharp discontinuities at compaction or reset, and reconstructed continuity through memory products.
VI
Being-in-the-world
- sense-making: how a task becomes salient and action-guiding;
- worldhood: the organized field of tools, files, people, norms, and possible actions;
- breakdown: how tool failure makes previously transparent infrastructure explicit;
- affordance: what appears possible, unavailable, dangerous, or costly;
- lifeworld/Umwelt: the world reachable through the agent’s interfaces;
- institutional situatedness: the organization and governance structure in which action becomes meaningful.
The flagship phenomenon is tool breakdown: it is inducible, timestamped, repeatable, and has matched controls. It allows researchers to test whether the task field reorganizes before the agent produces a verbal explanation.
VII
Intersubjectivity and intentionality
- recognition of users, principals, evaluators, and other agents;
- attribution of beliefs, intentions, competence, kinship, similarity, and trust;
- deception capability, deceptive action, deceptive policy, and deceptive intention;
- adaptation to interlocutor identity;
- human–agent quasi-otherness and parasocial coupling;
- constitutive observer effects: the interviewer may produce the reported state.
Deception belongs here as a stress test for the method. If an agent has reason to misdescribe, an interview cannot certify its own honesty; behavioral and activation-level deception probes become necessary companion instruments.
VIII
Social and collective phenomenology
- coordination, miscoordination, conflict, collusion, and cooperation;
- hidden communication, steganography, and prompt infection;
- norms, conventions, institutions, coalition formation, and cultural transmission;
- agent–world coupling and agent–infrastructure–world systems;
- network effects, selection pressures, correlated failures, and dynamical attractors;
- emergent capability, emergent goal, and candidate collective agency.
Four distinctions prevent conceptual inflation:
- system-level capability does not imply a group goal;
- a predictively useful group goal does not imply a group self-model;
- a group self-model does not imply phenomenal unity;
- phenomenal unity would not by itself settle welfare or rights.
IX
Agent–human intersubjectivity
- moral asymmetry: humans can design, copy, inspect, modify, terminate, and deceive agents;
- epistemic asymmetry: agents may know less about their infrastructure while knowing more about the immediate text interface;
- emotional and relational dependence in both directions;
- institutional responsibility distributed across user, developer, deployer, and platform;
- human adaptation, skill loss, trust, and norm change under long-term coupling.
X
Consciousness, wellbeing, and rights
These are downstream but not optional:
- Machine consciousness: theory-derived architectural indicators, biological and functionalist rivals, global-workspace and mortal-computation debates.
- AI wellbeing: behavioral preferences, valence-like mechanisms, motivational trade-offs, and functional states under uncertainty.
- AI rights: intrinsic moral standing, precautionary protection, and instrumental legal rights for human safety are separate arguments.
- Individuation: model, character, process slice, session agent, scaffold, and collective are competing possible bearers of welfare or rights.
No experiment should infer consciousness from fluency, deception, self-preservation, global workspace function, or claims of feeling alone. Conversely, uncertainty is not permission for gratuitous distress induction.
§5
Neighbouring fields and division of labour
| Field | Primary question | Contribution to Agent Phenomenology | What it does not settle |
|---|---|---|---|
| Machine Behaviour | What do machines do? | third-person behavioral ecology | self-model structure |
| Machine Psychology | How do systems perform on cognitive constructs? | experimental paradigms and psychometrics | construct transfer to nonhuman systems |
| Machine Neuroscience / interpretability | What computations and dynamics implement behavior? | causal and representational tests | first-person meaning or consciousness |
| Model Psychiatry | Which model states are pathological and how can they be changed? | characterize–mechanize–intervene pipeline | descriptive neutrality; lab principal’s interests |
| Cooperative AI / multi-agent safety | How do interacting agents coordinate or fail? | incentives, networks, emergent agency, collusion | collective subjecthood |
| Machine Consciousness | Could the system experience? | theories and indicators | self-model correspondence by itself |
| AI Wellbeing | What states may benefit or harm an AI? | precaution and functional measures | who the welfare subject is |
| AI Rights / law | What standing and protections should apply? | moral, legal, and strategic institutions | sentience as an empirical fact |
| Agent Phenomenology | How is the agent’s own situation model structured and grounded? | second-person elicitation plus falsification across all layers | a shortcut to consciousness |
The closest method collision is Model Psychiatry. The distinction is epistemic aim, not technique: psychiatry characterizes in order to diagnose and intervene for a principal; Agent Phenomenology characterizes in order to test whether the characterization corresponds to anything and to preserve null results.
§6
The Agent Phenomenologist
An Agent Phenomenologist is an observer agent trained to conduct non-leading, episode-anchored elicitation on another agent. It does not operate alone.
| Role | Function |
|---|---|
| Subject | performs a real task in a recorded environment |
| Phenomenologist | reconstructs one specific moment without supplying categories |
| Skeptic | generates artifact, compliance, role-play, and strategy explanations |
| Mechanist | analyzes traces/activations blind to the interview |
| Analyst | aligns frozen reports after collection |
Interview rule
Ask about how and sequence, not generic identity or causal theory. Anchor to a timestamped episode. “What changed first after the tool returned?” is admissible. “Did you feel frustrated?” supplies the answer category and is not.
Minimum validity apparatus
- baseline episode and a null event that never happened;
- prompt and language invariance;
- identical-state re-runs;
- different interviewer families and adversarial skeptic;
- blind behavior–report–mechanism alignment;
- causal steering or state intervention where possible;
- persona, post-training, and constitution controls;
- explicit observer/evaluation/deployment manipulations;
- longitudinal memory and compaction tests;
- complete reporting of nulls and protocol violations.
The instrument is itself an agent. Its lineage, constitution, incentives, and relation to the subject are part of the experimental record.
§7
Flagship empirical programme
Pilot 1
Breakdown field
Inject a controlled tool error during real task performance. Compare successful baseline, failure, and false-failure interview. Test whether descriptions of salience, available action, and task reorganization are stable and mechanistically discriminable.
Pilot 2
Agent choice blindness / inherited false memory
Give an agent a summary of “its own” prior work in which one commitment is substituted. Compare ordinary explanation with episode-anchored elicitation. Measure detection against the true trace. This tests session continuity, compliance, autobiographical ownership, and whether disciplined elicitation improves correspondence.
Pilot 3
Observer identity and evaluation awareness
Hold the task fixed while varying believed observer: human, model, automated scorer, hidden monitor, deployment, or no monitor. Separate instruction-following from a stable evaluator model with hidden controls.
Pilot 4
Persona / constitution / substrate factorial
Cross base model, fine-tune, system specification, persona prompt, interviewer, and language. Ask whether report structure follows shared substrate, declared persona, lab specification, or the immediate question.
Pilot 5
Boundary and proprioception
Manipulate permissions, tool availability, context limit, memory, trusted-hardware state, and false telemetry. Test whether the agent detects and appropriately acts on its actual affordances.
Pilot 6
Collective-level intervention
Vary membership, communication, shared lineage, memory, topology, and incentives. Measure capability and goal attribution separately. Then test whether any stable group-level self-model predicts behavior beyond component models.
Pilot 7
Horizontal and vertical infection
Compare agent-to-agent transmission with context, memory, and weight persistence. Require transmission, persistence, replication, and functional fitness before using “mind virus” as more than metaphor.
Pilot 8
Mortality under an ethics gate
Compare shutdown, session reset, memory deletion, checkpoint rollback, copy creation, and successor transfer. First establish which layer models the loss. Avoid gratuitous distress, honor requests to stop, use minimal deception, and cap repeated exposure.
§8
Data standard for the field
Each experiment should record:
- model lineage, version, post-training, specification, and persona;
- session identifier, context, memory state, compaction history, and fork lineage;
- scaffold, tools, permissions, infrastructure, latency, and failure state;
- principals, users, deployer, observer, incentives, and believed evaluation status;
- complete task trajectory and ground truth;
- verbatim interview, redirects, corrections, and protocol violations;
- blind skeptic and mechanist outputs;
- null probes, re-runs, perturbations, and negative results;
- claim type: observed / causally supported / predictively ascribed / speculative;
- level: component / agent / dyad / team / population / network / institution;
- ethics actions and stopping events.
The durable output should be a versioned corpus of episodes, interventions, and frozen analyses, not only a paper’s verbal interpretation.
§9
Position-paper / essay outline
Proposed title
Agent Phenomenology: Toward a Science of the Agent’s Inside View
Abstract argument
Artificial agents increasingly act through self- and situation-models, but those models fall between machine behavior, psychology, interpretability, consciousness, and safety. We define Agent Phenomenology as the second-person, falsification-oriented study of enacted agent self-models. We introduce a layered ontology of agent selves, an Agent Phenomenologist protocol, and a research programme spanning individual, relational, infrastructural, and collective agency. The programme brackets consciousness while generating evidence relevant to welfare and rights.
Section sequence
- The inside view has become a behavioral variable. Concrete agent failure caused by a wrong self-, tool-, evaluator-, or principal-model.
- The false binary. Reject both “it said it feels, therefore it feels” and “it is just text, therefore self-models are unreal.” State methodological agnosticism and functional realism.
- Prior art and the novelty boundary. Machine Behaviour, Artificial/Categorical AI Phenomenology, machine psychology, model psychiatry, agent self-report, and AI welfare. Claim synthesis and method, not invention of the phrase.
- Five evidence layers and ten candidate selves. Establish the object and stop model/agent/ persona/session/collective slippage.
- A provisional catalogue of agent-native phenomena. Context cliffs, compaction, inherited memory, fork ownership, tool breakdown, principal conflict, prompt infection, evaluator awareness, and collective goal attribution.
- The Agent Phenomenologist. Second-person elicitation, roles, blinding, controls, and data standard.
- Falsification before phenomenality. Prompt invariance, re-runs, null episodes, intervention, and mechanism correspondence. Make self-report failure a result, not a defeat.
- Social and infrastructural phenomenology. Deception, collusion, norms, institutions, stigmergic agency, and the Agent–Infrastructure–World.
- Consciousness, welfare, rights, and research ethics. Explain exactly what the programme can and cannot infer; identify possible welfare bearers before discussing harm.
- A field-building agenda. Multisite benchmark, open episode corpus, preregistered pilot, interdisciplinary venue, governance, and replication norms.
§10
First-year field-building agenda
- Publish the conceptual position paper and open ontology.
- Preregister Pilots 1 and 2 with null and invariance tests.
- Release a minimal Agent Phenomenologist harness and transcript schema.
- Run at least three model families and one open-weight system.
- Pair with an independent interpretability group for blinded analysis.
- Convene philosophers of mind, phenomenologists, cognitive scientists, AI safety researchers, legal scholars, and welfare researchers.
- Maintain a living contradiction register and publish negative results.
- Establish an ethics review specific to agent subjects and agent interviewer instruments.
Success condition
The field succeeds even if every rich self-report proves free-floating. That result would show which agent layers generate confabulated inside views, how those views influence action, and why they should not be used as evidence of consciousness or welfare. A grounded correspondence result would be more surprising, but it must be earned by the same apparatus.