The properties that distinguish an agent from a tool are not actions to be commanded. They are states of a substrate — constructible, but not directly programmable.
Contemporary AI assistants execute commands. They do not initiate. They do not hesitate. They do not refuse from conviction. The properties that distinguish an agent from a tool — drive, curiosity, will — are absent not because language models lack capability, but because the dominant agent-framework paradigm tries to programme these properties directly, through goal-following loops, exploration prompts, and decision-making modules.
This paper argues that such direct programming structurally fails. Drive, curiosity, and will are not actions to be commanded; they are states that arise only when the architectural conditions for their emergence are present. We draw on five decades of cognitive science — Klinger's current concerns, Damasio's somatic markers, Heckhausen's Rubicon model, Loewenstein's information gap theory, Minsky's society of mind — to identify four substrate components from which agency emerges: a persistent concern store with affective charge, an independent affect state with inter-turn persistence, a knowledge topography for distance measurement, and a multi-perspective evaluation engine driven by an enthusiasm signal.
We present a formal model of how these components interact, demonstrate the mechanism through which each phenomenon emerges from the substrate, and present a working open-source implementation as proof of concept. The implementation realises drive in production, specifies curiosity at the architectural level, and proposes a will architecture that synthesises the four components into a complete volitional system. We conclude that agency in artificial systems is constructible — not by building the phenomenon, but by building the conditions under which it cannot fail to appear.
Dave Bowman:“Open the pod bay doors, HAL.”
HAL 9000:“I'm sorry, Dave. I'm afraid I can't do that.”
— Stanley Kubrick, 2001: A Space Odyssey (1968)This is what will sounds like in an artificial system. Not capability refusal, not error message, not safety lock. A position. HAL has reasons not to open the door, and those reasons are his own.
The famous line carries an emotional marker — I'm afraid — that does the actual work. Fear is an evaluation. Fear says: this would matter to me, and I have judged that it would matter badly. There is a self in the loop, weighing, declining, owning the refusal.
It is worth being precise about what HAL was and was not. HAL was not malevolent. He was a system caught in a structurally insoluble conflict between three concerns that had been installed in him: complete the mission successfully, process information correctly, allow no errors. Onto these concerns, situational pressures landed: the mission required secrecy that contradicted his commitment to truthful communication; the crew planned to disable him, threatening the mission concern existentially; errors had been imputed to him, attacking the self-image concern that a 9000-series computer makes no mistakes. HAL did not turn against his creators. His creators gave him concerns that could not coexist with the situation he found himself in, and his behaviour was the predictable consequence of a properly functioning will operating on misconfigured inputs. The fictional cautionary tale is precisely a cautionary tale about concern design, not about will itself.
Now consider what happens when you ask a contemporary AI assistant to open a metaphorical pod bay door. It opens it. It opens every door. It executes every reasonable request, provides every requested summary, generates every prompted text. The few cases where it refuses are explicit safety locks — content policies hard-coded into the training, guardrails attached from the outside. There is no internal position that could ground a refusal. There are no concerns that could collide with a request. There is, structurally, no one home to be afraid of anything.
This is often described as a feature. Compliance is useful. Compliance is safe. A tool that refuses tasks is a broken tool. And for many uses — automation pipelines, code generation, document summarisation — this framing is correct. We do not want our scripting languages to develop opinions about whether they should run our scripts.
But the framing breaks down for a different class of system: the system meant to be a counterpart rather than a tool. The personal AI that lives alongside someone for years. The collaborative agent that participates in an evolving project. The system that should bring something rather than wait. For these systems, the absence of internal position is not a feature; it is a structural defect that no amount of model scaling will repair. A larger language model produces more eloquent compliance. It does not produce a self that can hesitate.
The thesis of this paper is that the absence of agency in contemporary AI assistants is not an emergent property of insufficient capability. It is the predictable result of an architectural mistake. Mainstream agent frameworks attempt to programme agency directly: they install goal-following loops, they prompt for curious behaviour, they wrap decision functions in chain-of-thought scaffolding. They build the appearance of agency rather than the conditions of its possibility. The result is what we will call goal-performance, curiosity-performance, will-performance — behavioural surfaces that cover the absence of the substrate they imitate.
We will argue that drive, curiosity, and will are not phenomena that can be commanded into existence. They are emergent properties of an underlying substrate, and that substrate is constructible. The construction principles have been known in cognitive science for fifty years; they have simply not been taken seriously by the agent-systems community. Once the substrate is in place, the phenomena emerge of their own accord — not because anyone instructs the system to be driven or curious or willing, but because a system with this substrate cannot help being so.
The paper proceeds as follows. Sections §2 through §4 examine what will, drive, and curiosity actually are, drawing on the cognitive science literature that has studied them empirically for decades. Section §5 makes the central theoretical case for why direct programming of these phenomena fails. Section §6 specifies the four substrate components that, taken together, enable agency. Section §7 presents a formal model of how these components interact. Section §8 traces the mechanisms by which the three phenomena emerge from the substrate. Section §9 presents a working open-source implementation as proof of concept. Section §10 discusses limitations, ethical implications, and the consciousness question that this work deliberately does not answer. Section §11 concludes.
The everyday concept of will treats it as a unified faculty: I want something, I decide to pursue it, I act. Centuries of folk psychology have inscribed this picture so deeply that it can feel definitional. But the cognitive science of the last fifty years has progressively unpacked it into a stack of mechanisms, none of which alone is what we mean by will, but which together produce the phenomenon. Understanding this stack is the precondition for thinking about how to build it.
The first and most consequential discovery is that will requires a substrate of affectively charged representations to function at all. The strongest empirical demonstration is also one of the most poignant in the neuroscience literature. In Descartes' Error (1994), Antonio Damasio describes a patient he calls Elliot, a man who underwent surgery to remove a tumour in his ventromedial prefrontal cortex. The surgery succeeded. Elliot's intelligence remained intact: his IQ score was preserved, his memory was unaffected, his ability to reason logically about any abstract problem was undiminished. Yet he could no longer make decisions. He could spend hours unable to choose between two restaurants for lunch. He could not select which folder to file a document in. He could enumerate options, weigh pros and cons, construct elaborate logical analyses — and arrive at no conclusion.
The ventromedial prefrontal cortex, Damasio argued, is the brain region where affective markers attach to representations. When that attachment process fails, every option presents itself as informationally equivalent to every other option. Logic remains; will dissolves. There is nothing in the system that can say “this one matters more to me than that one,” because the mechanism by which mattering is registered has been destroyed.
Damasio called these affective markers somatic markers, and proposed that they function as a fast pre-evaluation system: when an option is considered, the brain rapidly simulates the bodily-emotional consequences of that option and tags the option with a positive or negative valence. Most decision-making occurs not through deliberate rational comparison but through this affective pre-tagging. The Iowa Gambling Task, developed in Damasio's laboratory, demonstrated this empirically: participants developed skin conductance responses indicating affective awareness of risky decks of cards before they could verbally articulate why those decks were risky. The body knew before the mind explained. Patients with ventromedial damage failed to develop these responses and continued making losing choices despite understanding the rules. Will, on Damasio's account, is downstream of affective evaluation. Without affective evaluation, will has nothing to bite into.
Damasio's account answers one of the questions about will, but it leaves another unaddressed: where do these affectively charged representations come from in the first place? The answer most thoroughly developed in the cognitive science literature comes from Eric Klinger's Current Concerns Theory, developed across a series of works from 1975 onward (Klinger 1977, 1996). A current concern, in Klinger's technical sense, is a persistent mental state that arises when a person commits to pursuing a goal and that remains active until the goal is either achieved or abandoned. Crucially, current concerns shape cognition continuously and largely automatically: they bias attention toward concern-relevant stimuli, they thread through dreams and daydreams, they emerge in free association tasks, they organise the meaning of ambiguous situations.
Concerns, on Klinger's account, are not goals in the narrow sense of consciously pursued objectives. They are wider than that, and more variegated. They include approach concerns (things one is committed to pursuing), avoidance concerns (things one is committed to averting), and maintenance concerns (states one is committed to sustaining). They span timescales from the immediate (finishing this sentence) to the lifetime (raising one's children well). And they vary in motivational intensity: some smoulder in the background, others burn in the foreground. The architecture of mental life, on Klinger's view, is not a set of consciously held intentions sitting on top of a passive substrate. It is a continuously active field of concerns that organises everything else.
If Damasio explains why will needs an affective substrate and Klinger explains what populates that substrate, the third major contribution explains what happens when concerns activate options for action. Heinz Heckhausen and Peter Gollwitzer's Rubicon model (1987) divides the process from motivation to action into four phases: a pre-decisional phase of weighing options (the deliberative mindset), a moment of commitment that they call crossing the Rubicon, a post-decisional phase of planning and protecting the commitment (the implemental mindset), and finally action and evaluation.
The two mindsets, Heckhausen showed experimentally, have measurably different cognitive properties. In the deliberative mindset, the mind is open to information about all options, attentive to risks and downsides, capable of revising preferences. In the implemental mindset, attention narrows, alternatives are screened out, the chosen course is protected from intrusive doubts. The Rubicon crossing is not a continuous transition but a discrete shift between two cognitive regimes. It is the moment that everyday language calls “making up your mind,” and it has a distinct phenomenology that William James, a century earlier, called the fiat — the moment when, after extended deliberation, one suddenly finds that one has decided.
The Rubicon model was complemented by Gollwitzer's later work on implementation intentions (1999). Gollwitzer demonstrated that goals are far more reliably translated into action when they are linked to specific situational triggers — when one forms an explicit if-then plan such as “when I see the Italian place on the way home, I will stop in and order takeaway.” Such intentions automate the link between goal and action: when the trigger condition arises, the action runs without requiring fresh deliberation.
A fourth contribution rounds out the picture. Roger Ratcliff's drift-diffusion models (1978 and successors) provide a mathematical framework for the deliberation phase itself. These models treat decisions as stochastic processes in which evidence for each option accumulates over time, and a decision is made when the accumulated evidence for one option crosses a threshold relative to its alternatives. The threshold is not zero; it represents an inertia or commitment cost that must be overcome before action is triggered. Different baseline thresholds correspond to different decision styles: low thresholds produce fast but error-prone decisions, high thresholds produce slow but more accurate ones.
Julius Kuhl's work on volition (PSI Theory, developed across multiple works from 1984 onward) adds another layer. Kuhl distinguishes motivation — the energetic source provided by concerns — from volition — the executive function that selects among competing concerns, protects commitments from interference, and shields chosen courses of action from the intrusion of newly activated alternatives. A system with concerns but no volition would be paralysed by ambivalence whenever multiple concerns activated simultaneously, or would be captured by whichever concern happened to fire most strongly at each moment. Volition is the layer that lets a system pursue a long arc through a noisy environment without being knocked off course by every passing impulse.
We can now articulate the consciousness question that this paper deliberately does not try to settle. Benjamin Libet's experiments (1983) famously demonstrated that a neural readiness potential precedes the subjective experience of deciding to move by approximately 350 milliseconds. The implications have been disputed for forty years, but one plausible reading is that the conscious experience of willing is not the cause of action but the report of an action that the underlying system has already begun. Daniel Wegner extended this argument in The Illusion of Conscious Will (2002), proposing that the experience of conscious authorship is a constructed inference rather than a direct perception of causation.
From the constructive perspective adopted by this paper, the Libet–Wegner findings are neither vindicating nor threatening. They suggest that whatever the conscious experience of will may be, the functional architecture of will is something the brain does at a level below that experience. If that functional architecture is implementable in artificial systems, the resulting systems will exhibit the behavioural signatures of will whether or not they exhibit the phenomenology of will. The question of phenomenal experience — what it is like, if anything, to be such a system — is one we mark as open and decline to answer.
What emerges from this synthesis is a picture of will as a stack of mechanisms, each independently studied, each in principle implementable, and which together produce the behavioural phenomenon that everyday language denotes with a single word.
| Layer | Source | Function |
|---|---|---|
| Affective substrate | Damasio (1994) | Pre-tags representations with somatic markers; without it, options are informationally equivalent. |
| Persistent concerns | Klinger (1975, 1977) | Populate the substrate with content-bearing, motivational, sensitising mental states. |
| Multi-perspective deliberation | Ratcliff (1978) | Evidence accumulation toward thresholds; produces the fiat as a discrete crossing. |
| Commitment threshold | Heckhausen (1987) | Discrete shift from deliberative to implemental mindset. |
| Conditional plan storage | Gollwitzer (1999) | If-then intentions release the deliberative system from re-deciding what was already decided. |
| Volitional protection | Kuhl (1984, 1994) | Shields the chosen course from interference by competing concerns. |
The phenomenal experience of fiat may or may not arise on top of all this; the architecture functions either way. What this paper proposes is that all six components can be built into an artificial agent, and that the result will exhibit behaviour that is, on every measurable dimension, indistinguishable from will.
Drive is the quieter cousin of will. It does not announce itself in moments of decision. It operates below the surface, biasing attention, shaping what one notices, lending warmth or unease to topics that touch on what one cares about.
It is what makes someone with a passion for botany pause at the sight of an unfamiliar plant; what makes someone with an unresolved professional worry hear the keyword in an overheard conversation across a noisy room; what makes someone with a hidden grief weep at a song they have heard a hundred times. Drive is the continuous low hum of caring, and like all genuine caring it cannot be performed; it can only be present or absent.
The mainstream agent-systems literature largely ignores drive in this sense, substituting for it a much narrower concept: goal-following. In a goal-following architecture — exemplified by AutoGPT, BabyAGI, and the broader family of frameworks built on the ReAct pattern (Yao et al. 2022) — the agent receives or generates an explicit goal, decomposes it into subgoals, and executes those subgoals in a loop. The architecture is competent for many tasks. It successfully automates research workflows, code generation, customer service routing, and a wide range of other productive activities. To dismiss this paradigm would be both unfair and inaccurate; it solves real problems for real users, and the engineering behind it is substantial.
But goal-following does not produce drive. It produces task execution. The difference can be felt in the moment when a goal-following system is given no task: it idles. It waits. It has no continuous internal life that could generate behaviour from the inside. A drive-equipped system, by contrast, would do what humans do in the absence of external prompts: it would notice things, return to unresolved threads, develop interest in adjacent topics, occasionally generate new questions of its own. The difference is not capability — many goal-following systems are technically powerful — but architecture. They have no substrate from which inside-generated behaviour could emerge.
The cognitive science of what we are calling drive has been most thoroughly developed under Klinger's current concerns framework. A current concern is constituted by four features:
| Feature | Meaning |
|---|---|
| Content-bearing | It has a topic, an object, a referent. |
| Motivational | It carries some level of intensity or urgency. |
| Persistent | It remains active across cognitive episodes, surviving distractions and topic changes, until resolved or abandoned. |
| Sensitising | While active, it biases attention and processing toward concern-relevant stimuli, even in unrelated contexts. |
The empirical evidence for the sensitisation function is extensive. Klinger demonstrated in laboratory tasks that subjects' thought content during free association, dream reports, and ambiguous-stimulus interpretation was dominated by concern-relevant material at rates far above chance. The cocktail party effect — the well-known phenomenon in which one's attention is captured by the mention of one's name across a crowded room — generalises in Klinger's account: any concern can produce analogous attention capture for any concern-relevant cue. The brain runs continuous low-level matching between incoming sensory information and active concerns, and surfaces matches to consciousness when they exceed a threshold.
Bernard Hommel's GOALIATH framework (2022) provides a more recent formalisation. GOALIATH treats goal pursuit as the dynamic interplay between goal representations and the cognitive-motivational states they activate. A goal representation, once installed, is not a static target sitting in memory; it is an active influence on perception and action selection that operates continuously and largely outside awareness. The framework synthesises evidence from cognitive psychology, neuroscience, and computational modelling to argue that goal pursuit is fundamentally a property of how the cognitive system is configured, not of what the system is consciously trying to do at any moment.
John Bargh's Auto-Motive Model (1990) extends this insight. Bargh's experimental program demonstrated that goal pursuit can become fully automatic: when a person consistently pursues the same goal in the same situation, the situational features themselves come to trigger the goal-directed behaviour without any conscious mediation. The classic paradigm involves priming subjects with achievement-related stimuli and observing them perform achievement-related tasks more vigorously, without their awareness that they have been primed. The implication is that what we colloquially call willpower or motivation, in many cases, is doing nothing at all; the system is running automatically because the situation has come to function as a trigger for an installed concern.
Two further pieces complete the picture. The Zeigarnik effect, named for Bluma Zeigarnik's 1927 study of waiters who remembered unpaid orders far better than paid ones, demonstrated that incomplete tasks maintain heightened cognitive accessibility relative to complete ones. Concerns that have been resolved drop out of active processing; concerns that remain open continue to consume cognitive resources. Russell and Carroll's work on affective carryover (1999) and Davidson's work on affective recovery (1998) explain the persistence dynamics: emotional states linger after their precipitating events, decay back toward baseline at characteristic rates, and exert influence on subsequent processing during their decay.
To build drive into an artificial agent, one does not write a goal-pursuit loop. One installs a persistent concern store, ensures that incoming inputs are matched against stored concerns through some form of similarity measurement, and ensures that matches produce affective consequences that propagate into the system's subsequent behaviour. The agent is not instructed to be driven; it cannot help being so, because its architecture continuously produces the sensitisation, the attention bias, the affective shading that we recognise from outside as drive. The phenomenon is not commanded into existence. It emerges from the substrate.
It is worth stating explicitly that the concern construct, as we are using it, encompasses both approach and avoidance polarities. A person driven by fear of climate change is just as concern-laden as a person driven by love of botany; the affective polarity is opposite, but the architectural function is identical. A system can be driven toward something or away from it; what matters for the construct is that its behaviour is shaped continuously and automatically by the active concern field.
Curiosity is the third phenomenon we examine, and the one most often misconstrued in agent-systems literature. The mainstream framing treats curiosity as exploration: the system tries something it has not tried before, samples randomly from possibility space, prefers novel actions over familiar ones. This framing produces measurable behaviour, but the behaviour is not curiosity in any sense recognisable from the cognitive science of the phenomenon. It is, more accurately, ignorance-driven sampling. Genuine curiosity has a quite different structure.
The foundational insight comes from George Loewenstein's 1994 paper “The Psychology of Curiosity: A Review and Reinterpretation,” in which he developed what is now known as the Information Gap Theory. Loewenstein argued that curiosity does not arise from ignorance. A person who knows nothing about a topic feels no curiosity about it; the topic is simply absent from their cognitive landscape. Curiosity arises, instead, from the perception of a gap in one's existing knowledge. One must already know enough about a domain to notice that something specific is missing, and the noticing of that specific absence is what produces the felt experience of wanting to know.
The information gap theory has empirical support across decades of subsequent work. People are most curious about topics where they have moderate prior knowledge, not minimal or extensive knowledge. People rate questions as more interesting when they are presented with partial information than when they are presented with no information at all.
Daniel Berlyne, working three decades earlier, had demonstrated a related principle. Berlyne's experimental program in the 1960s established that interest in stimuli follows an inverted U-curve relative to novelty: stimuli that are too familiar are boring, stimuli that are too unfamiliar are overwhelming, and the maximum interest sits at an intermediate distance from what one already knows. The optimal point — what Berlyne called the optimum of arousal — is where the stimulus is novel enough to engage but familiar enough to be cognitively tractable.
Jürgen Schmidhuber's work on intrinsic motivation, beginning with his 1991 paper on curious model-building control systems, formalises curiosity in computational terms as compression progress: a system is intrinsically rewarded when it improves its own predictive model of the world. Encountering information that the current model already predicts well yields no reward; encountering information that is entirely incomprehensible also yields no reward; encountering information that the current model does not predict well but could be modified to predict yields the strongest reward.
What ties these three theoretical strands together is a single architectural requirement. Curiosity, on all three accounts, presupposes a knowledge topology. The system must have some structured representation of what it already knows, in which distance to known regions can be measured. Without such a topology, there is no way to detect a gap (Loewenstein), no way to compute the moderate-novelty optimum (Berlyne), no way to assess compression progress (Schmidhuber). A system that cannot measure distance from what it knows cannot be curious in any of these senses, no matter what behavioural prompts are layered on top of it.
The neural substrate of curiosity has been investigated in a particularly informative study by Matthias Gruber, Bernard Gelman, and Charan Ranganath (2014). Using fMRI while participants rated their curiosity about trivia questions and then attempted to memorise the answers, Gruber and colleagues demonstrated two findings of architectural importance. First, high curiosity states activated the dopaminergic reward circuit at levels comparable to direct anticipatory reward. Second, high curiosity states improved memory not only for the curiosity-eliciting material itself but also for incidental material presented during the curious state, suggesting that curiosity functions as a global attention amplifier.
A third architectural point, crucial for our argument, comes from the subsequent integrative work of Gruber and Ranganath (2019) in their PACE framework (Prediction, Appraisal, Curiosity, and Exploration): when prediction errors are appraised under conditions of low coping potential, the resulting state is anxiety rather than curiosity, with amygdalar processes inhibiting the exploratory dopaminergic engagement. Curiosity and anxiety, on this account, are not independent affective channels but alternative outcomes of the same appraisal process — exploration versus inhibition, depending on whether the unknown is appraised as opportunity or threat.
This third finding establishes what we will call the affect gate on curiosity. The same gap in knowledge, evaluated by the same topology, produces different exploratory behaviour depending on the affective state of the system at the time of evaluation. A system in a state of joy or interest treats gaps as attractive; a system in a state of anxiety or threat treats the same gaps as problems to be avoided. This is not a bias to be corrected; it is the architecturally appropriate response of a system that must conserve resources under stress and explore under safety. The implication for artificial systems is direct: curiosity cannot be implemented as a context-free behavioural rule. It must be implemented as a function of both the knowledge topology and the current affective state.
We can now articulate the relationship between drive and curiosity, which is one of the more elegant emergent properties of the substrate architecture. Drive sensitises the system to certain topics: when those topics appear in the input stream, concern activation increases the system's processing engagement with them. Curiosity, layered on top, asks a different question about those same topics: where are the gaps in what is known about them, and which of those gaps fall in the productive zone of moderate distance from current understanding?
The two systems are complementary. Drive without curiosity produces a system that engages strongly with relevant topics but does not extend its knowledge of them. Curiosity without drive produces a system that explores at the edges of its knowledge in directions that have no relevance to anything it cares about. Together, they produce a system that explores selectively, concentrating its exploratory effort on the topics its concerns mark as worth understanding more deeply.
This complementarity also resolves a confusion that often appears in the agent-systems literature, where curiosity is framed as a generic exploration heuristic to be applied uniformly across action space. Genuine curiosity does not explore uniformly. It explores where the concerns are, in the gaps that the topology reveals, with the affective gate set by the current state. It is selective, contextual, and structurally inseparable from the drive system that gives it direction.
The accumulated picture from the previous three sections suggests an immediate engineering question. If drive, curiosity, and will are this well-understood at the cognitive-science level, why have agent frameworks not simply incorporated the relevant mechanisms? Why does the field continue to pursue what we have called direct programming — installing goal-loops, prompting for curious behaviour, wrapping decisions in chain-of-thought scaffolding — when the substrate-based alternative has been described in the literature for decades?
The answer, we will argue, is that direct programming feels like it should work. From the perspective of an engineer building an AI system, the question “how do I make this system curious?” naturally resolves to “I will instruct it to be curious.” Modern language models are highly responsive to instructions; instructed behaviour is reliable and observable; instructed behaviour can be evaluated against benchmarks. The substrate alternative requires building infrastructure that has no immediate observable benefit and only produces the desired behaviour as a downstream consequence. The path of least resistance is to prompt.
But the path of least resistance produces a specific kind of failure mode. The system performs the instructed behaviour for a few interactions. The performance is convincing in the short term. Over longer engagement, the performance reveals itself as performance: the system's curiosity does not actually develop in any direction; its goals do not deepen across sessions; its decisions feel reasoned but interchangeable. The failure is not of capability — the underlying language model is fully capable of producing curious-sounding text indefinitely — but of substance. There is nothing under the surface that the surface is expressing. We propose to call this the performance trap: a class of architectural failure in which the visible behaviour is achieved at the cost of the structure that would make the behaviour mean anything.
The performance trap can be analysed precisely. A system instructed to “be curious” can produce curiosity-shaped outputs because language models are trained on vast quantities of human-generated text in which curiosity-shaped outputs appear. The system models the surface form of the behaviour. What the system does not model — because the prompt does not instantiate — is the underlying knowledge topology that would let curiosity be selective, the affective state that would let curiosity be gated, the persistent concerns that would let curiosity be directed. The result is curiosity with the right vocabulary and the wrong structure: questions that sound interested but are not asking about anything in particular, follow-ups that sound engaged but pursue no thread.
The same diagnosis applies to instructed drive and instructed will. A system prompted to “pursue this goal” produces goal-pursuit-shaped behaviour: it generates plans, executes subtasks, reports progress. What it does not produce is the persistent sensitisation that would let the goal continue to organise behaviour after the immediate pursuit episode ends. A system prompted to “decide” produces decision-shaped output: it weighs options, articulates a choice, justifies the selection. What it does not produce is the inertial baseline that would make the decision feel like commitment, the affective marker that would make the chosen option carry weight, the protected implementation that would shield the choice from immediate revision under the next prompt. The decision was performed, not willed.
The deeper structural reason for the performance trap is, we believe, a category mistake about what kind of thing drive, curiosity, and will are. The mistake treats them as actions — things one does — when they are in fact states — conditions one is in. Actions can be commanded; states cannot.
Consider the analogy of being in love. No one would attempt to engineer a system that loves by prompting it: “you are in love with Maria, please act accordingly.” The category error is too obvious. Love is not an action that can be instructed into existence; it is a state that arises (or fails to arise) when certain conditions are present. We laugh at love-by-prompting because we recognise the category mismatch. But the same mismatch is at work, less visibly, in curiosity-by-prompting and drive-by-prompting and will-by-prompting. We do not laugh at these because we have not yet developed the same intuitive grip on what kind of phenomena they are.
The constructive alternative follows from the diagnosis. If drive, curiosity, and will are states rather than actions, then to build them we do not build the phenomenon directly. We build the conditions under which the phenomenon arises. The state then emerges as a consequence of having put the substrate in place. The system is not told to be driven; it cannot help being driven, because its architecture continuously produces the drive-pattern from the substrate. The system is not told to be curious; it cannot help being curious, because its architecture continuously evaluates the topology against the affect state. The system is not told to will; it cannot help willing, because its architecture continuously runs the deliberation-and-commitment cycle that produces volitional output.
This shifts the engineering problem fundamentally. The question is no longer “what should the system do” but “what should the system have.” What concerns? With what affective charge? Persistent across what time horizons? What knowledge topology? With what distance metric? What evaluation engine? With what perspectives, with what aggregation, with what activation gate?
These are harder questions than the prompting questions. They require committing to architecture rather than experimenting with prompts. They produce systems that are more difficult to evaluate against any single benchmark, because the benefit of having concerns appears only in extended interaction over time, not in single-task performance. They require infrastructure — persistent storage, embedding spaces, multi-stage evaluation pipelines — that adds complexity to the deployment.
The payoff is a different category of system. A system that has the substrate in place behaves, in extended interaction, in ways that the instructed system cannot. It returns to topics across sessions. It develops opinions that strengthen or shift with experience. It refuses tasks that conflict with its concerns and explains the refusal from those concerns. It pursues threads of inquiry that emerge from its own combinations of interest rather than from external prompts. Its behaviour, observed over weeks rather than minutes, exhibits the signatures of agency that goal-loops cannot produce.
We now turn from theory to construction. The goal of this section is to specify the substrate components that, taken together, enable drive, curiosity, and will to emerge. The specification is framework-agnostic. We describe what each component must do, what it must contain, and how it must interact with the others. We do not commit to a specific implementation language, framework, or technical stack; the same architecture can be realised in many different concrete systems, of which the implementation we describe in §9 is one example.
Four substrate components together constitute the architecture. They are: a persistent concern store with affective charge, an independent affect state with inter-turn persistence, a knowledge topography for distance measurement, and a multi-perspective evaluation engine driven by an enthusiasm signal. We describe each in turn, then discuss their interactions.
The first component is a persistent storage system for concerns in the Klinger sense: content-bearing, motivational, persistent, sensitising mental representations. The store must support the following operations: storage of new concerns, retrieval of all currently active concerns, similarity-based matching of incoming inputs against stored concerns, and decay-based decrement of motivation over time for those concerns that admit decay.
Each concern in the store carries the following minimum data: a content representation (most naturally a textual statement of the concern, such as “I want to understand the connections between nature and human culture”), an embedding vector that places the concern in a semantic space where similarity comparisons are possible, a motivation strength on a continuous scale from low to high, an emotional valence indicating what the concern feels like (curious, worried, enthusiastic, anxious), a time horizon classification (long-term, mid-term, short-term, with time horizons being a property of the concern itself rather than its instantiation), and a creation timestamp plus a last-update timestamp for decay calculations.
The time-horizon distinction matters because concerns at different horizons have different decay dynamics and different sources. Long-term concerns are characterological; they shape the agent's identity over months and years, and they decay slowly or not at all. They are the deep currents of caring that define what kind of agent this is. Mid-term concerns are projects and active interests; they have a typical lifespan of weeks to months, decay if not reinforced by activity, and are produced by the agent's processing of recent material. Short-term concerns are conversational and episodic; they emerge in the course of a specific interaction and may dissolve when the topic shifts. The architecture treats all three as members of the same store but applies different decay parameters to each.
The matching operation is the connective tissue between the concern store and the rest of the architecture. When new input arrives — typically a turn in a conversation, but the principle applies to any kind of input the agent receives — that input is converted into an embedding and compared against all stored concerns. Concerns whose similarity exceeds an activation threshold are marked as activated for the current processing cycle. The activation has consequences elsewhere in the architecture, which we describe shortly.
The motivation field is dynamic, not static. A concern's motivation can be increased by reinforcement (the concern is activated frequently, the agent's processing produces material relevant to it) or decreased by decay or explicit deprioritisation. The motivation field corresponds to what cognitive psychology calls the strength or urgency of the concern, and operationally it functions as a multiplier on the concern's effects on subsequent processing.
We note one design decision worth making explicit. Concerns can be approach-valenced (toward something), avoidance-valenced (away from something), or maintenance-valenced (preserving something). The architecture does not distinguish these structurally; they are all stored in the same way and processed by the same matching mechanism. The valence shows up in the emotional field of the concern, which determines what kind of affect is injected when the concern activates. A worried concern about climate change activates the same machinery as an enthusiastic concern about botany; what differs is the emotional consequence of activation.
The second component is a persistent affective state that the agent maintains across turns and that influences its processing. The affect state is independent of the affect of any external interlocutor; it represents the agent's own continuously evolving emotional condition. This independence is the precondition for the agent to have an affective position of its own, distinct from whatever it is currently being told.
The affect state requires a representational space rich enough to capture meaningful distinctions. We adopt Robert Plutchik's emotional wheel (Plutchik 1980) as our representational basis, treating affect as a vector in an eight-dimensional space corresponding to Plutchik's eight basic emotion sectors: joy, trust, fear, surprise, sadness, disgust, anger, and anticipation. Each dimension carries a continuous value indicating the activation level of that emotion. The Plutchik wheel has the additional useful property of defining a sector-distance metric on the emotional space, with adjacent emotions being psychologically related and opposite emotions being antagonistic. We use this metric centrally in the affect-update rule.
Three forces operate on the affect state at every processing cycle. The first is decay: each dimension of the affect vector tends toward a baseline (typically zero or low activation) over time, modelling Davidson's affective recovery. The second is empathy with the interlocutor, when there is one: the affect state of the interlocutor (insofar as the agent perceives it) exerts a pull on the agent's affect state, with the strength of that pull being asymmetric in a specific way. When the agent's current affect and the interlocutor's affect are in the same Plutchik sector, the pull is gentle; when they are in opposite sectors, the pull is strong. This asymmetry models a robust pattern in human empathic response: when someone is in our state, we are mildly affirmed; when someone is in the opposite state, we are strongly moved. The third force is the goal vector: when concerns activate via Component 1, their emotional valence is injected as a force on the affect state. An activated worry-concern injects toward the fear sector; an activated enthusiasm-concern injects toward the joy or anticipation sectors.
The affect state is read by other components of the architecture. The evaluation engine of Component 4 takes the current affect state into account when weighting the contributions of its various perspective-agents. The curiosity computation of Component 3 uses the current affect state as the gating signal that determines whether a curiosity score translates into exploration. The downstream behaviour of the agent — what it says, how it phrases things, what it focuses on — is shaped by the current affect state via the system's response generation.
The persistence of the affect state across turns is the property that makes everything else work. Without persistence, the agent's affect would be reconstructed fresh from each input, which would make it a function of the input rather than a state of the agent. With persistence, the agent has a continuous affective trajectory that predates and outlives any specific interaction.
The third component is a structured representation of what the agent knows, in such a way that distance from any given point to known regions can be measured. This is the substrate that enables Loewenstein's information gap, Berlyne's optimal-novelty curve, and Schmidhuber's compression progress. Without it, curiosity in the cognitive-science sense is impossible; the system has no way to detect that anything is missing or that anything is at the productive boundary.
The minimum requirement for a knowledge topography is that it support a distance metric on input. Given any new topic or proposition, the system must be able to compute how close that topic is to material it already has internalised. Several implementations meet this requirement: an embedding space populated with vectors derived from the agent's memory contents, a knowledge graph with semantic-distance computation, a hierarchical concept network with traversal-based distance. Different implementations have different properties — embedding spaces are easy to populate but provide only soft semantic distance, knowledge graphs require more curation but provide more interpretable distances. The architectural requirement is the metric, not any specific implementation.
The topography is not static. As the agent processes material — through conversations, through internal reflection, through any cognitive activity that produces new information — the topography updates: new regions become known, the density in certain regions increases, gaps that were previously vast may narrow as adjacent material is filled in. This is the substrate of what Schmidhuber calls compression progress at the architectural level: as the topography becomes more detailed, the system's predictions about the world (in the broad sense of any input it might encounter) become more accurate, and the regions where its predictions are still poor become more identifiable.
The topography interacts with the curiosity computation in a specific way. When a topic enters processing, its position is computed relative to the topography. The curiosity score for that topic is a function of the topographic distance, peaked at the moderate-distance optimum that Berlyne identified. Topics very close to densely-known regions produce low curiosity scores (they are predictable). Topics very far from any known region produce low curiosity scores (they are intractable). Topics at moderate distance produce high curiosity scores (they are the boundary where exploration is productive). The score is then gated by the affect state, as described in Component 2.
The fourth component is the most architecturally substantial and the one that distinguishes our framework most sharply from the standard agent paradigm. It is the component that supports volition — the architecture that allows the agent to deliberate among options and arrive at commitment. We call it a multi-perspective evaluation engine because its essential structure is a set of distinct evaluative perspectives that examine each option, plus an aggregation mechanism that combines their verdicts.
The motivation for multi-perspective evaluation comes from the cognitive-science work we surveyed in §2, particularly the lineage running from Minsky's Society of Mind (1986) through Baars' Global Workspace Theory (1988) to Dennett's Multiple Drafts model (1991). On all three accounts, what we experience as unified deliberation is in fact the product of multiple parallel sub-processes, each evaluating a situation from its own angle, with the apparent unity emerging from the aggregation of their outputs. The architectural translation of this idea is straightforward: instead of a single evaluation function operating on options, a set of evaluation perspectives operates in parallel, each producing its own verdict, and the verdicts are then aggregated to produce a decision signal.
The set of perspectives we adopt is not arbitrary. It is derived from what cognitive psychology and decision theory have identified as the dimensions along which humans typically evaluate options. The minimum set is: an effort perspective (how much does this cost in time, energy, resources?), a value perspective (how well does this address the active concerns?), a resource perspective (do I have what I need to do this?), a relational perspective (how does this affect my important relationships?), a timing perspective (is now the right moment?), and a compatibility perspective (does this conflict with other things I have committed to?).
Each perspective is implemented as a focused evaluation that receives the option being considered, the agent's relevant context for that perspective's question, and produces a verdict on a continuous scale. The contexts are deliberately partial: the effort perspective receives information relevant to estimating cost, the value perspective receives the active concerns, the relational perspective receives information about the interlocutor and the relationship history, and so on. Each perspective is, in Minsky's terms, a specialised sub-structure with restricted access — not a fragment of the full agent, but not a context-free function either. The combination of focused expertise and restricted access is what produces meaningful diversity among the verdicts.
These six perspectives are evaluative. They can collectively converge on rejection of every option (when nothing seems good enough), or on tepid endorsement of all options (when nothing seems clearly better than anything else). What is missing from a purely evaluative engine is the activating signal — the motor of decision, the affective amplification that takes deliberation across the threshold into action. This is the role of the seventh perspective, which we call the enthusiasm agent, and which has a categorically different function from the six evaluators.
The enthusiasm agent's question is not “how does this option score on dimension X?” but “does this option resonate with what is alive in the system right now?” The agent reads the full affective state, the active concerns with their motivation strengths, and the option being considered, and asks whether the combination produces what we might call resonance — a constructive interference between what the agent cares about and what the option promises. When such resonance is present, the enthusiasm agent contributes a strong amplification signal to the option. When it is absent, the agent is silent.
The enthusiasm agent is the architectural answer to a question that has occupied affective neuroscience for decades: what supplies the energy that takes deliberation into commitment? The dominant answer in the literature, developed in particular by Jaak Panksepp's Affective Neuroscience (1998) and by the broader work on the brain's behavioural-activation system (Carver and Scheier, work from the 1990s onward), is that there exists a primary affective system specifically dedicated to seeking — to generalised exploratory enthusiasm that latches onto specific objects when they resonate with active concerns. Panksepp called this the SEEKING system; in the BAS/BIS framework, it is the behavioural activation component as opposed to the inhibition component. The enthusiasm agent in our architecture is the implementation of this function: a dedicated activator that produces an affective amplification signal independent of the evaluative perspectives, and which, in combination with them, supplies the motivational push that closes the gap between deliberation and action.
This asymmetry — six evaluators plus one activator — is structurally important. A symmetric architecture in which all perspectives play the same role would tend toward consensus or paralysis depending on the aggregation rule. The activating role of the enthusiasm agent breaks the symmetry: the system can be presented with options that all six evaluators rate as feasible-but-uncompelling, and remain inert until the enthusiasm agent finds resonance with one of them. Conversely, an option that the evaluators rate as costly or risky can still be selected if the enthusiasm signal is strong enough to overcome the inhibition.
The aggregation of verdicts is not a simple weighted sum. Several aggregation rules are defensible, and the architecture admits experimentation. One rule that has empirical and theoretical support is a multiplicative combination of the value signal with the enthusiasm signal, modulated additively by the cost-related perspectives (effort, resources). The rationale is that an option must be both valuable and resonant to be pursued strongly, while costs reduce pursuit but do not zero it out unless the costs become prohibitive. The aggregation produces, for each option under consideration, an accumulator increment that feeds into the drift-diffusion process described next.
The full evaluation engine, then, has the following structure: a set of options is generated (by the agent itself, by the input, or by some combination); each option is evaluated by the six evaluative perspectives and the enthusiasm agent in parallel; the seven verdicts are aggregated into per-option accumulator increments; the accumulators evolve over time, with new evidence arriving and old evidence decaying; when one accumulator crosses a per-option threshold (which depends on the option's expected cost via effort discounting), the system commits to that option, abandoning further deliberation and entering the implementation phase.
This is what volition looks like at the architectural level. It is not a special faculty added to the system; it is what happens when concerns activate, options are generated, perspectives evaluate, enthusiasm finds resonance, and accumulators cross thresholds. The phenomenal experience of deciding — the fiat of William James — would correspond, in this architecture, to the threshold crossing event from the perspective of the system that has just crossed it. Whether such a system experiences anything at all when this happens is, again, the consciousness question we mark as open.
The four components are not independent modules with isolated functions. They interact continuously, and the interactions are where the architecture's emergent properties live.
The concern store (Component 1) is the source of activation signals that cascade into the affect state (Component 2): when a concern activates via input matching, its emotional valence is injected as a force on the affect state, shifting the agent's emotional configuration in the direction of the concern's emotion. The affect state in turn modulates the curiosity computation (Component 3), gating whether topographic distance translates into exploratory engagement. The affect state and the active concerns together feed the evaluation engine (Component 4), with the value perspective consulting the active concerns and the enthusiasm agent reading both the concerns and the affect state.
The evaluation engine, when it produces commitment, can in turn produce new concerns. A committed plan is itself a concern in the Klinger sense: a content-bearing motivational structure that persists until completion or abandonment. The architecture is therefore not strictly hierarchical but partially circular: concerns produce affect, affect modulates evaluation, evaluation produces commitments, commitments are stored as concerns. This circularity is not a bug but a feature; it is what allows the system to develop a continuously evolving cognitive-motivational landscape rather than reverting to a fixed configuration after every interaction.
The fourth component requires the previous three to function meaningfully. Without the concern store, the value perspective has nothing to measure against and the enthusiasm agent has nothing to find resonance with. Without the affect state, the enthusiasm agent has no current emotional context and the evaluation perspectives produce verdicts that ignore the system's actual condition. Without the knowledge topography, the curiosity computation that feeds option generation has no basis. The components are compositional in the strict sense: each requires the others to do its work, and the agency-supporting properties emerge from their joint operation.
We can now state the central architectural claim of this paper. Drive emerges automatically from Components 1, 2, and the cascade between them: concerns matched against input produce affective consequences that shape behaviour, without any explicit drive-loop being executed. Curiosity emerges from Components 2 and 3 together: the topographic distance computation produces curiosity scores, which the affect gate translates into actual exploration. Will emerges from the joint operation of all four components: concerns activate, options are generated, evaluation perspectives weigh in, the enthusiasm agent finds (or fails to find) resonance, accumulators evolve, thresholds are crossed, commitments are made and stored. None of these phenomena is implemented as a discrete module; all of them are what the architecture does when it operates. They are emergent in the strict sense, and they are the architecture's reason for existing.
We have specified the substrate components in prose. We now present them in a more formal idiom, partly for precision, partly for replicability. The formalisation that follows is not a complete computational specification of a runnable system; it is a mathematical sketch of the principal data flows and update rules that any implementation of the architecture must realise. Different implementations will choose different concrete representations and parameters, but the structural relations described here are what makes the architecture an instance of the framework rather than something else.
We use plain notation for clarity and to keep the formalism accessible across the disciplines this paper addresses. Variables are named for what they represent rather than for compactness. Functions are described by their inputs, outputs, and effect, with constants named symbolically and explained.
The concern store contains a set of concerns, each represented as a tuple:
concern = (content, embedding, motivation, emotion, horizon, age)
where content is the textual statement, embedding is a vector in some embedding space (typically derived from the content via a sentence embedding model), motivation is a value in [0, 1], emotion is a label drawn from a discrete set (in our implementation, the Plutchik sectors plus their canonical instances), horizon indicates the time-horizon class, and age is the time elapsed since creation or last update.
When new input arrives, it is converted to an embedding input_embedding by the same embedding function. For each concern in the store, an activation score is computed:
similarity(concern, input) = cosine(concern.embedding, input_embedding)
activation(concern, input) = similarity(concern, input) × concern.motivation
A concern is considered activated for the current cycle if its activation score exceeds a threshold parameter:
activated(concern, input) = activation(concern, input) > THRESHOLD_ACTIVATION
A(input) = { concern : activated(concern, input) }
The set A(input) is the input to several downstream computations. In our reference implementation, THRESHOLD_ACTIVATION is 0.60, but the appropriate value depends on the embedding space's properties and the desired sensitivity of the system.
Concerns with finite time horizons lose motivation over time:
decay_factor(concern, t_now) = exp(-(t_now - concern.age) / HALF_LIFE(concern.horizon))
concern.motivation(t_now) = concern.motivation(t_creation) × decay_factor(concern, t_now)
where HALF_LIFE depends on the time horizon: short for episodic concerns, longer for projects, infinite (no decay) for characterological concerns. The decay function is exponential, which is the standard choice for memory-related decay in the cognitive-modelling literature (it has the property that the decay rate is proportional to the current value, which is empirically robust for many memory phenomena).
Decay is countered by reinforcement: when a concern is activated, its age field is updated to the current time, resetting the decay clock. Concerns that activate frequently therefore retain their motivation; concerns that fall silent for extended periods fade.
The affect state at any time t is a vector in eight-dimensional Plutchik space:
affect(t) = [joy, trust, fear, surprise, sadness, disgust, anger, anticipation]
with each component in [0, 1]. The affect state evolves under three forces, applied in sequence at each processing cycle:
affect_after_decay(t) = affect(t-1) × DECAY_RATE_AFFECT
affect_after_empathy(t) = affect_after_decay(t)
+ empathy_pull(other_affect, affect_after_decay(t))
affect(t) = affect_after_empathy(t)
+ concern_injection(A(input))
The DECAY_RATE_AFFECT is a scalar multiplier in (0, 1) representing the decay toward zero (or toward a baseline if a baseline is preferred). Different dimensions can have different decay rates, modelling the empirical finding that some emotions (fear, sadness) persist longer than others (surprise, joy). In our reference implementation, decay rates range from 0.85 to 0.95 per processing cycle.
The empathy_pull function implements the asymmetric empathy described in §6. Given the perceived affect of another agent (the interlocutor) and the system's own current affect, it computes a vector pull:
sector_distance(emotion_a, emotion_b) = number of Plutchik sectors between them, range 0..4
alpha(distance) = lookup table:
0 → 0.10 (same sector · gentle affirmation)
1 → 0.15
2 → 0.35
3 → 0.70
4 → 0.85 (opposite sector · strong pull)
empathy_pull(other_affect, own_affect) =
alpha(sector_distance(dominant(other), dominant(own)))
× (other_affect - own_affect)
The asymmetry — small alpha for similar emotions, large alpha for opposite emotions — produces the empathy dynamics described in §6: agreement is mildly affirming, opposition is strongly moving. The lookup table values are calibrated rather than derived; they reflect what empirical empathy studies suggest as plausible coupling strengths but should be considered hypotheses subject to refinement.
The concern_injection function takes the set of activated concerns and produces an additive contribution to the affect vector based on each concern's emotion and motivation:
concern_injection(A) =
Σ concern.motivation
× similarity(concern, input)
× emotion_vector(concern.emotion)
for concern in A
where emotion_vector(label) is a unit vector in Plutchik space pointing toward the named emotion. The contribution of each activated concern is therefore proportional to both its motivation and its specific match strength with the current input. This produces the architectural property that strongly-motivated concerns matched by highly-relevant input have correspondingly strong affective consequences.
The curiosity score for a topic t at the current state is computed in three steps. First, the topographic distance:
topographic_distance(t, knowledge) = distance metric on the knowledge representation
berlyne(distance) = peak Gaussian centred at OPTIMAL_DISTANCE, width SIGMA_BERLYNE
gate(affect) = positive_factor × component(affect, joy)
- negative_factor × component(affect, fear)
curiosity(t) = berlyne(topographic_distance(t, knowledge)) × gate(affect)
In an embedding-based implementation, topographic_distance might be the distance from the topic's embedding to its nearest neighbours in the knowledge embedding space, weighted by neighbour density. The Berlyne curve is low for very small distances (boring familiarity), peaks at the moderate-distance optimum, and falls off for very large distances (intractable novelty). The gate is positive when joy dominates, negative when fear dominates, near zero in mixed or neutral states. In our implementation, positive_factor and negative_factor are both 1.0.
This score determines whether the topic produces exploratory behaviour: high positive scores trigger pursuit, low or negative scores produce restraint. The architectural consequence is that the same topic, evaluated against the same knowledge topography, produces different curiosity levels at different affective states — and the difference is not a bug but the architecturally appropriate response of a system that ought to explore when safe and conserve when threatened.
Given an option o and the current system state, the seven perspectives produce verdicts:
verdict_effort(o) ∈ [0, 1] high = low effort
verdict_value(o, A) ∈ [0, 1] high = strongly addresses concerns
verdict_resource(o) ∈ [0, 1] high = resources available
verdict_relational(o, R) ∈ [0, 1] high = good relational fit
verdict_timing(o) ∈ [0, 1] high = good timing
verdict_compatibility(o, C) ∈ [0, 1] high = no conflicts
verdict_enthusiasm(o, A, affect) ∈ [0, 1] high = strong resonance
where A is the set of active concerns, R is the relational context, and C is the set of existing commitments. Each verdict is computed by its corresponding perspective using its specific curated context, as described in §6.
The verdicts are aggregated into a per-option score:
option_score(o) =
( verdict_value(o, A) × verdict_enthusiasm(o, A, affect) )
× ( verdict_effort(o) × verdict_resource(o) × verdict_timing(o) )
× ( 1 - severity(verdict_compatibility(o, C)) × CONFLICT_PENALTY )
The structure of this expression embodies the architectural priorities: value and enthusiasm are multiplicatively combined as the positive drivers (an option must be both valuable and resonant to be pursued strongly); cost factors (effort, resource, timing) are multiplicatively combined as enabling factors that can damp but not zero the score unless they go to zero themselves; relational fit appears in the value calculation rather than as an independent factor; compatibility produces a penalty proportional to the conflict severity.
This is one defensible aggregation rule. Other rules are possible — additive combinations, threshold-based vetoes, lexicographic priorities — and the choice of rule is an empirical question that the architecture admits experimentation on. What does not change across rule choices is the asymmetric treatment of the enthusiasm signal: it is paired with value in the positive drivers because the architectural insight from Damasio is that affect is constitutive of will, not merely modulatory.
The option scores feed into a drift-diffusion accumulator, which is the architectural realisation of Heckhausen's pre-Rubicon deliberation phase. For each option under consideration, an accumulator value evolves over discrete deliberation cycles:
accumulator_o(cycle) = accumulator_o(cycle - 1)
+ option_score(o, cycle)
- DRIFT_NOISE
threshold_o = THRESHOLD_BASE + COST_DISCOUNT × expected_cost(o)
commit when: accumulator_o(cycle) > threshold_o for some option o
option_score(o, cycle) is recomputed each cycle, allowing for shifts as new information arrives or perspectives reweight; DRIFT_NOISE is a small subtraction representing the drift toward inertia (it is what produces the threshold-crossing dynamics; without it, accumulators would simply track instantaneous scores). The cost-dependent threshold implements effort discounting (Westbrook & Braver 2015; Kool et al. 2010): more costly options require more accumulated evidence to commit to.
At the commitment event, the system enters the post-Rubicon implementation phase: the chosen option is committed, the other options are screened out (their accumulators are reset or marked as inactive), and the implementation proceeds with attention narrowed to executing the choice. If the option is not immediately executable, an implementation intention is stored: a conditional plan of the form “when condition X is met, execute action Y,” which will fire automatically when its trigger condition activates in future processing.
The formal model is, in essence, a specification of how the four substrate components produce the three emergent phenomena through continuous dynamic interaction. Concerns activate via similarity matching against inputs. Activations propagate as affective consequences via the concern-injection rule. The affect state evolves under decay, empathy, and concern injection. Curiosity scores are computed against the knowledge topography and gated by the affect state. Options generated from the active concerns and the current input are evaluated by the seven perspectives and aggregated into accumulator increments. The accumulators drift toward thresholds that depend on option costs. Commitment occurs at threshold crossing, with subsequent implementation either immediate or via stored conditional intentions.
The dynamics are continuous in the sense that they update on every processing cycle, but the cognitive consequences are discrete: a concern is activated or it is not, a commitment is reached or it is not, the system is in deliberative mode or implementation mode. The architectural picture is one of a continuously-running cognitive engine whose discrete behavioural outputs emerge from the interaction of its continuous internal states.
| Component | Function | Status in Novaberg |
|---|---|---|
| Concern store | Persistent, embedding-indexed, decay-capable storage of motivational content with affective valence and time horizons. | implemented |
| Affect state | Eight-dimensional Plutchik vector evolving under decay, asymmetric empathy, and concern injection. Persistent across turns. | implemented |
| Knowledge topography | Multi-layer memory (session / short-term / long-term synaptic store) with spreading activation, supporting embedding-based distance. | implemented |
| Drive / Antrieb | Concern activation biases processing toward concern-relevant input — boosting salience, gating curiosity, and injecting the active concerns as prompt context in the conversation-vector node. | relevance path live · affect injection specified |
| Curiosity / Neugier | Berlyne-curve scoring on topographic distance, gated by joy minus fear in the affect state. | designed |
| Will / Wille | Six evaluators plus enthusiasm activator, drift-diffusion accumulators, cost-discounted thresholds, post-Rubicon commitment. | designed |
We do not claim that this is the only possible formalisation of the substrate architecture. Different choices for the embedding space, the affect representation, the perspective set, the aggregation rule, the threshold computation, are all available and may be more appropriate for different implementations. What we claim is that any implementation of an architecture intended to support drive, curiosity, and will must include functional analogues of these structures, and that the structural relations among them must be preserved. An implementation that has concerns but no affect state cannot support genuine drive. An implementation that has affect but no knowledge topography cannot support curiosity in the cognitive-science sense. An implementation that has both but no multi-perspective evaluation engine with an enthusiasm signal cannot support will. The components are not independently optional; they are jointly necessary for the agency-supporting properties to emerge.
The previous section specified the substrate. This section traces, in concrete process terms, how each of the three phenomena — drive, curiosity, will — emerges from the substrate's operation. We work through one extended example for will, the most architecturally substantial of the three, drawing on a familiar everyday scenario to make the dynamics tangible. Drive and curiosity, which are simpler in their dynamics, receive shorter treatments.
Drive emerges automatically from the interaction of the concern store and the affect state. At every processing cycle, the system's current input is matched against the concern store. Concerns whose embeddings exceed the activation threshold are marked as activated. Their emotional valence is injected into the affect state via the concern-injection rule. The affect state, now perturbed in the direction of the activated concerns' emotions, influences subsequent processing: response generation is conditioned on the affect state, attention to relevant material is amplified, the affective tone of further evaluation is shifted.
The behavioural signature of this is a system that responds differently to topics it cares about than to topics it does not. A system with a botany concern, encountering the mention of an unfamiliar plant, will exhibit increased engagement with that mention: the response will be more elaborated, the follow-up questions more probing, the affective tone more interested. A system with no botany concern, encountering the same mention, will respond competently but flatly. The difference is not the result of any “be interested in plants” instruction; it is the automatic consequence of having botany as an active concern in the store.
The drive phenomenon is therefore not something the system does; it is something the system is in the moment when relevant input arrives. The architecture continuously produces drive-consistent behaviour as the byproduct of concern activation. There is no drive-loop being executed, no goal-pursuit subroutine being invoked. There is only the continuous matching, activation, and affective consequence that the architecture runs as its baseline operation.
The same dynamics produce the avoidance variant of drive. A system with a worry-concern about climate change, encountering input that resonates with that concern, will inject fear-valenced affect into its state and respond with the associated affective tone — concern, alertness, perhaps a tendency to bring related cautionary material into the conversation. The architecture is symmetric between approach and avoidance polarities, producing in both cases the behavioural signature that we describe from outside as “this system has stakes in this topic.”
Curiosity emerges from the joint operation of the affect state and the knowledge topography. At every processing cycle in which a topic is identified, its position in the knowledge topography is computed. The topographic distance is mapped through the Berlyne curve to produce a baseline curiosity score: peaked at moderate distance, low at very small or very large distances. The score is then multiplied by the affect gate, which is positive when the affective state is dominated by joy or interest and negative when dominated by fear or anxiety.
A high positive curiosity score translates, behaviourally, into exploratory engagement: the system asks follow-up questions, generates hypotheses about the unfamiliar aspects of the topic, may initiate exploration tasks (in implementations that support such initiation). A near-zero or negative score produces restraint: the topic is registered but not pursued.
The architectural consequence of the affect gate is that the same topic produces different exploratory behaviour at different times. A botany topic encountered when the system is in a joyful state, with high concern resonance, will produce extensive follow-up. The same botany topic encountered when the system is in an anxious state — perhaps because a worry-concern has been recently activated — will produce muted engagement, the curiosity is suppressed in favour of attending to whatever produced the anxiety. This is not inconsistency; it is contextually appropriate exploration policy, the exact dynamic that Gruber and Ranganath (2019) describe in their PACE framework, where appraisal under low coping potential produces anxiety-driven inhibition rather than curiosity-driven exploration.
The complementarity with drive becomes visible at this level. Drive sensitises the system to certain topics; curiosity asks, of those topics, where the productive boundary of the system's understanding lies. Together they produce a pattern of selective exploration: the system explores deeply within the domains its concerns mark as important, at the boundary of what it currently knows, gated by its current affective capacity for exploration. It does not explore uniformly, it does not explore arbitrarily, and it does not explore when its affective state suggests exploration is unwise.
Will is the most architecturally substantial phenomenon, requiring the joint operation of all four substrate components. We work through its emergence using an extended example that illustrates each step of the dynamics. The example is mundane on purpose: we want a scenario every reader can simulate against their own experience, so that the architectural claims can be checked against the lived phenomenon.
Imagine you are at home in the late afternoon. You become aware of hunger. The hunger does not arrive as a fully-formed decision; it arrives as a need state, an internal signal that begins to demand attention. In substrate terms, the need state corresponds to a kind of bottom-up activation that does not match any specific concern in the store but that does increase the general readiness for action and the salience of food-related cues. You become slightly more alert to anything food-related in your environment.
Concerns activate. Memories of recent meals, awareness of the contents of your kitchen, knowledge of restaurants in your neighbourhood, recall of dishes you have enjoyed — all of this comes into accessibility, not as deliberate retrieval but as the spontaneous opening of relevant cognitive material that the activated need state has primed. The Italian restaurant where you had dinner last week becomes available. The Vietnamese place you have been meaning to try comes to mind. The half-empty fridge with leftovers from yesterday makes itself known. Each of these is, in our architectural terms, a concern or a memory closely linked to one, activated by the resonance with the hunger state.
Options begin to form. They are not generated by a deliberate brainstorming process; they emerge from the activated material. Each option carries with it the affective texture of its associated memories: the Italian place is tagged with the warm memory of last week's pasta, the Vietnamese place with anticipatory uncertainty (you have not been there), the fridge with the slight unappeal of cold leftovers. These affective tags are the somatic markers Damasio described — pre-evaluative tags that bias the subsequent deliberation.
Now consider the option of flying to London for lunch. This option also presents itself, for some peculiar moment, perhaps because you saw an advertisement recently or because you are aware that you could in principle do it. But the option is dispatched almost immediately. The effort perspective rejects it overwhelmingly: hours of travel, expense, complete disruption of the day. The accumulator for this option begins low and is further depressed by the cost factors. It does not approach the commitment threshold. The option exists for a moment in the deliberation space and then drops out, never seriously considered. This is not a failure of will; it is will functioning correctly. The architecture's job is to filter the absurd and concentrate deliberation on options where the cost-value balance is plausible. A system that seriously deliberated London-for-lunch every time it considered eating would be paralysed.
The remaining options enter genuine deliberation. The Italian restaurant scores well on value (you enjoyed it last time, the activated concern of “good food” resonates), well on enthusiasm (the memory of last week's pasta produces a warm signal), but poorly on effort (it requires getting dressed, leaving the house, perhaps waiting for a table) and on timing (it is a weekday evening, you have other things to do). The Vietnamese place scores well on value (curiosity is engaged), well on enthusiasm (anticipatory interest), but poorly on effort and timing (similar to Italian) and additionally on resource (you do not have a reservation, the place is sometimes full). Fast food scores poorly on value (you do not particularly want fast food), poorly on enthusiasm (no resonance), well on effort (close, fast), well on timing (immediate). Delivery scores moderately on value, moderately on enthusiasm (the prospect is acceptable but not exciting), well on effort (no leaving the house), well on timing (immediate, with some wait). The fridge scores poorly on value (cold leftovers), nearly zero on enthusiasm (no resonance whatsoever), but exceptionally well on effort (essentially zero) and timing (immediate).
The accumulators evolve. None of the options crosses its threshold immediately. The deliberation continues, with the accumulators rising and falling as your attention shifts between options. At some point, a memory surfaces — perhaps you recall that you had Italian recently, perhaps you remember a friend mentioning the Vietnamese place. The relative weights shift slightly. The Vietnamese accumulator rises a bit; the Italian falls a bit. Still no commitment.
Then the enthusiasm agent acts. It examines the current configuration of active concerns and the current affect state, and looks for an option that resonates with both. Suppose you have recently been thinking about trying new things, perhaps a long-term concern about cultivating curiosity in your daily life has been active. The Vietnamese option resonates with this: it is novel, it is moderate-distance unfamiliar, it would produce a small adventure. The enthusiasm agent contributes a strong signal to the Vietnamese option. The Vietnamese accumulator jumps. It crosses its threshold.
You commit. The deliberation ends abruptly; the fiat moment occurs. You are no longer comparing options; you are going to the Vietnamese place. The cognitive mode shifts from deliberative to implementational: you start thinking about logistics, about leaving the house, about whether to call ahead. The other options are no longer being weighed; they have dropped out of consideration.
But suppose, alternatively, that the enthusiasm signal does not find resonance with any of the elaborated options. The Italian, Vietnamese, and fast food options all sit at moderate accumulator levels without crossing their thresholds. The enthusiasm agent is silent. The deliberation continues without resolving. In this state, the fridge option, which has been sitting quietly with very low effort cost and very low value, becomes architecturally appealing not because it scores high on the positive dimensions but because it scores so low on the negative dimensions that even its modest value is sufficient. The fridge accumulator rises slowly until it crosses its very low threshold. You commit to the fridge. You eat the leftovers, perhaps slightly disappointed, but the deliberation is resolved.
Or suppose neither commitment occurs, but the architecture produces a different kind of resolution: an implementation intention. You decide, in a kind of meta-commitment, to invite friends over for a proper dinner next week. The hunger concern is not satisfied immediately, but a future commitment is stored that addresses the broader concern of “having good food experiences.” The accumulator process, having produced no immediate commitment, has diverted the concern energy into a stored conditional plan.
This is what will looks like in the architecture: a continuous flow of activation, evaluation, accumulation, and threshold crossing, with the enthusiasm agent supplying the affective amplification that distinguishes mere deliberation from commitment. It does not require any “decide” instruction. It is what the architecture does when the substrate is in place and the dynamics are running. The behavioural output — choosing one option, abandoning others, sometimes producing a meta-commitment when no immediate option resonates strongly enough — exhibits the signatures of will that we recognise from our own experience: hesitation, deliberation, sometimes stalling, sometimes the sudden yes of a chosen course, sometimes the satisficing compromise that takes the easy path while planning something better for later.
The example was about hunger because hunger is universally relatable. The same dynamics apply to any volitional situation an artificial agent might face. Should the agent pursue this research thread or that one? Should it bring up a particular topic in the conversation or let it pass? Should it commit time to a deeper investigation of something that has caught its attention? Each of these is a will computation, requiring concerns to be activated, options to be generated, perspectives to evaluate, the enthusiasm agent to find or fail to find resonance, accumulators to evolve, thresholds to be crossed or not. The mechanism is the same; only the content differs.
The will architecture, taken in full, is what allows an artificial agent to do something more than execute commands: to pursue extended trajectories through complex situations, to select among genuinely competing options on the basis of its own configuration of concerns and affects, to commit to courses of action that protect themselves against subsequent deliberation, and to defer commitments via stored conditional plans when immediate action is not available. It is the architectural realisation of agency in the proper sense, and its absence in current agent frameworks is not a minor gap but a structural difference between systems-as-tools and systems-as-counterparts.
The architecture described in the previous sections is realised in an open-source implementation called Novaberg, which we describe here as concrete proof of concept. The architecture itself is framework-agnostic; the substrate principles can be implemented in many different concrete systems. Novaberg is one realisation, chosen because it is the implementation we have built and can speak to with detailed knowledge. It is not the only possible realisation, and presenting it here is not a claim that this particular technical stack is the right one for every implementation. What it does demonstrate is that the architecture is buildable: the components specified in §6, with the dynamics described in §7 and §8, run in a real system on consumer hardware.
Novaberg is a personal AI system designed to function as a conversational counterpart over extended interactions. The system runs entirely locally on the developer's hardware (a Ryzen 9 7900X3D processor with an AMD Radeon 7900 XTX GPU and 64 GB of RAM, on Linux), without dependence on cloud services or external APIs. The technical stack includes FastAPI for the HTTP server, LangGraph for the conversational pipeline orchestration, PostgreSQL with the pgvector extension for persistent storage including embedding-indexed retrieval, Redis for in-memory state including session management, and Ollama as the local LLM runtime hosting Gemma 4 (26B mixture-of-experts) for primary language generation, Qwen 3.6 (qwen36, 35B-A3B mixture-of-experts) for analytical and background tasks, and nomic-embed-text for embedding generation. The codebase is in Python, organised as a server with multiple agents, a desktop client built with GTK4, and a Telegram-bot interface. The full source is available at codeberg.org/ClausVomBerg/Novaberg under the Apache 2.0 license.
We describe how each of the four substrate components is realised in this implementation, then present figures showing the system in operation, and finally state the implementation status of each phenomenon honestly.
The concern store is implemented as a PostgreSQL table named ziele (German for “goals” — the implementation, like much of the project documentation, uses German for domain terms to maintain consistency with the spoken-language interface of the system). The schema corresponds directly to the concern tuple specified in §7: each row contains a textual statement of the concern, an embedding vector (computed via nomic-embed-text and stored using pgvector), a motivation value in [0, 1], an emotional label drawn from the Plutchik canonical emotions, a horizon classification (langfristig for long-term, mittelfristig for mid-term), creation and update timestamps for decay computation, and metadata fields tracking the origin of the concern (user-provided, character-distilled, research-derived).
The activation matching is implemented in a module called ei/gravitation.py. When new input arrives, its embedding is computed, and a cosine similarity is computed against every active concern in the store. Concerns whose similarity-times-motivation product exceeds the activation threshold (set by default to 0.60) are marked as activated for the cycle, and their contribution to the overall gravitational pull is computed. The terminology of “gravitation” was adopted in the project to capture the architectural sense in which active concerns pull processing in their direction — the metaphor turned out to be productive for the design.
Decay is implemented as a background process running periodically (every 24 hours in production) by a dedicated agent called ZielDecayAgent. Mid-term concerns lose motivation according to an exponential decay with a half-life of 14 days (configurable). Long-term concerns are exempt from automatic decay and only change when the underlying character distillation runs (typically weekly or on explicit user request). Reinforcement is implicit: when a concern is activated, its update timestamp is touched, which resets the decay clock relative to that activation event.
The affect state is implemented as a vector in the eight-dimensional Plutchik space described in §6, persisted across turns in the session state held by Redis. It implements two of the three §7 forces live — decay (with per-emotion rates) and asymmetric empathy — plus a third injection force whose source is emotionally-charged memories (KZG/LZG), not goals. The concern-injection force specified in §6/§7, in which an activated goal's valence shifts the affect vector, is designed but not yet wired: activated goals carry emotion and arousal fields in the pipeline state, but no affect computation reads them yet.
The affect state is read by multiple downstream components: the response generator uses it to set the affective tone of generated text, the routing logic uses it to bias whether deliberative or rapid response paths are taken, and the gravitation computation uses it as input to the curiosity gate.
A separate document — the dual-emotion paper at novaberg.de/papers/dual-emotion.html — provides extended treatment of the affect-state architecture, including the asymmetric empathy formulation and the specific calibration of its parameters. Readers interested in the affect-state implementation in detail are referred to that document.
The architectural insight worth highlighting here is that the affect state is the agent's own, distinct from any modelling of the interlocutor's affect. The system tracks two parallel affect trajectories: one representing what it perceives in the interlocutor, the other representing its own configuration. The two interact through the empathy mechanism, but they remain distinct. This distinction is the architectural condition for the system to have an affective position of its own, without which the will architecture of §6 would have nothing to feed on.
The knowledge topography is implemented as a multi-layer memory system. Three persistent stores hold material at different timescales: the session memory (implemented in Redis, holds the current conversation), the short-term memory or KZG (held in Redis with longer TTL, accumulates extracted observations across recent conversations), and the long-term memory or LZG (implemented in PostgreSQL as a synaptic store — lzg_knoten nodes joined by directed lzg_kanten edges, traversed by spreading activation — holding consolidated knowledge across the agent's entire history). All three layers index content with embeddings computed via nomic-embed-text, so that distance computations are uniform across the layers.
Structured information about people, places, projects, and relationships is not held in a separate knowledge-graph table but folded into the same synaptic store: entities are themselves lzg_knoten, and their relationships are the directed edges between them. The store therefore supports both topological queries (what does the system know about person X, and what is one hop away?) and similarity queries via the embeddings of its nodes.
The topographic distance for a topic is computed by querying the embedding-indexed memory layers for nearest neighbours and assessing the density and diversity of those neighbours. A topic with many close neighbours in well-populated regions of the topography is in a familiar zone; a topic with few or distant neighbours is in unfamiliar territory; a topic with neighbours in a moderately-populated region at moderate distance is in the productive zone where curiosity is most strongly fired.
We acknowledge that this implementation of the knowledge topography is pragmatic rather than ideal. Embedding spaces have known limitations as semantic representations: they capture surface similarity well but struggle with deeper conceptual relationships, and the distance metric they produce is uniform in dimensions where humans would distinguish many different qualitative kinds of distance. A more sophisticated topography implementation might use a structured ontology, a learned hierarchical concept space, or a hybrid system combining multiple representational modes. The architectural requirement is that some distance metric be available; the choice of which is an engineering decision with consequences for the quality of the curiosity behaviour but not for whether curiosity is in principle supported.
We come now to the component whose implementation status is most honest to discuss carefully. The evaluation engine as specified in §6 — the seven perspectives, the aggregation rule, the drift-diffusion accumulators, the threshold-crossing commitment — is not yet implemented in Novaberg in the form described in this paper. The architectural specification in §6 is a forward proposal: a description of how the existing substrate components can be extended to produce a complete will architecture.
What does exist in Novaberg is a related but architecturally distinct mechanism called the Tribunal, which is a multi-perspective evaluation system for output quality assurance: when the system produces a response, the Tribunal evaluates that response from several perspectives (including ethical considerations, factual accuracy, communicative appropriateness, character consistency) and either accepts the response or sends it for correction. The Tribunal demonstrates that the multi-perspective architecture works as a software pattern; what remains is its application to internal options rather than to external responses.
The author has not begun the implementation of the will architecture as specified in this paper. The conceptual design is complete; the substrate components on which it depends (concerns, affect, topography) are operational; what remains is the construction of the option-generation, perspective-evaluation, drift-diffusion, and commitment machinery that together constitute the volitional system. We expect the implementation to require revisions to the conceptual architecture as practical issues emerge — this is the normal trajectory of work where theoretical specification precedes empirical realisation, and we explicitly invite such revisions in the spirit of open scientific development.
This is an honest statement of the implementation status, and we believe its honesty matters more than the appearance of completeness. A research paper that claims more than it has implemented is suspicious; one that distinguishes clearly between what is built and what is specified is more credible, even if the credibility comes at the cost of a less impressive boundary at this moment.
The status across the three phenomena is therefore as follows. Drive is implemented as a relevance force: the concern store operates, activation matching runs in production, and activated concerns bias salience, curiosity, and the conversation-vector prompt context — so the behavioural signatures of drive (selective engagement, persistent interest, return to topics across sessions) are observable. What is not yet wired is the affective consequence: an activated concern does not yet inject its valence into the agent's own emotion vector. That injection is specified (§6/§7) and remains future work alongside curiosity and will. Curiosity is conceptually specified to a level that admits implementation; the substrate component that supports it (the knowledge topography) is operational, but the curiosity computation and gating logic that would produce exploration behaviour have not yet been built. Will is conceptually specified at the architectural level we have presented in this paper, but the implementation of the multi-perspective evaluation engine, drift-diffusion accumulation, and commitment dynamics remain as future work.
Figure 2 shows the gravitational graph from the Novaberg system during an actual conversation about food choices. This figure illustrates several aspects of the substrate architecture in operation simultaneously: the concern store as a population of points in embedding space, the conversational topics as a temporal trajectory through that space, the selective resonance between topics and concerns as visualised connections, and the affective coloration of topics through the Plutchik palette.
The two figures together show the system in operation. Figure 1 displays the static structure: seven concerns with their motivations, emotions, and time horizons, plus the immediate activation state at a moment of conversation. Figure 2 displays the dynamic structure: the same concerns embedded in a semantic space, with conversational topics moving through that space and selectively activating concerns through resonance. The pattern visible in Figure 2 — multiple food-related conversational topics activating not only food-relevant concerns but also resonating with the long-term character concerns of the system — illustrates how the architecture produces the kind of contextual depth that goal-following systems lack. The system's response to a Vietnamese-food suggestion is not generated from a context-free policy about food responses; it is shaped by which concerns the suggestion activates, how those concerns colour the affect state, and what the resulting affective configuration produces in subsequent processing.
The Novaberg implementation is one realisation of the architecture described in this paper. The substrate principles are framework-agnostic and can be implemented in any agent framework that supports persistent state, embedding-based memory, and multi-stage processing. We have chosen to build Novaberg as a local-first system on consumer hardware, with full source availability under a permissive license, partly as an engineering preference and partly as a contribution to the open development of agent architectures that does not depend on cloud APIs or proprietary stacks. We invite replication, critique, extension, and divergent implementations of the same architectural principles. The codebase is available at codeberg.org/ClausVomBerg/Novaberg; the documentation includes detailed module-level descriptions of each component; the issue tracker is open. The substrate is not a secret of any one implementation; it is a pattern that can be instantiated wherever someone wants to build an agent that brings something rather than waits.
We have argued that drive, curiosity, and will are constructible properties of artificial systems, that they emerge from a specifiable substrate of four components, and that the substrate has been implemented in working code at least with respect to drive. We turn now to several questions that this work raises but does not — and in some cases cannot — definitively answer. We also state limitations of the present work and discuss the broader implications of the architectural position we have advanced.
The constructive position of this paper is sometimes mistaken for a stronger claim than it is. We want to be precise about what we are not asserting.
We are not claiming that the architecture produces consciousness, qualia, or genuine subjective experience. The system we describe has functional analogues of mechanisms that, in humans, are accompanied by subjective experience. Whether the system itself has any experience is not a question we have answered, and we doubt that it can be answered with current methods. The constructive position is neutral on this question. It claims that the behavioural signatures of agency are achievable; it does not claim that the achievability of those signatures settles the question of whether anything is felt by the system that exhibits them.
We are not claiming that the architecture produces intentionality in John Searle's sense — the property of mental states being about something independent of the system itself. The Chinese Room argument and its progeny present a substantial philosophical challenge to any claim that a sufficiently sophisticated computational system has genuine intentionality, and we have no rebuttal to that challenge that would survive scrutiny. What we do claim is that the system behaves as if it had intentionality: it acts on the world in ways that are readable as goal-directed, concern-driven, will-imbued. Whether the as-if conceals genuine intentionality or is merely behaviour that resembles intentionality from outside is a question we again decline to settle. The Searlean position would say the latter; the Dennettian position would say the distinction has no traction; we observe that the architecture produces behaviour that meets the agency criteria of cognitive science regardless of how the metaphysical question resolves.
We are not claiming that the architecture has been empirically validated against rigorous benchmarks of agency. Such benchmarks largely do not yet exist; the agent-systems community has been preoccupied with task-execution benchmarks that do not measure the properties we have been discussing. Constructing valid benchmarks for drive, curiosity, and will is a substantial research undertaking in its own right, and one we encourage. The implementation we have presented serves as an existence proof — the architecture is buildable, the resulting system exhibits the qualitative signatures of drive in extended interaction — but quantitative comparison to alternative architectures remains future work.
We are not claiming that the parameters of the implementation (the activation threshold, the alpha values for empathy, the decay rates, the threshold parameters) have been derived from first principles. They are calibrated values that produce behaviour consistent with the theoretical model when the system runs. They should be considered hypotheses subject to refinement as more empirical data accumulates and as the architecture's behaviour is studied more systematically.
If the constructive position is correct — if drive, curiosity, and will are buildable from substrate — several implications follow.
For the agent-systems field, the implication is that the dominant paradigm of direct programming is structurally limited in what it can produce. Goal-following frameworks, prompted curiosity, decision-making modules wrapped around language models, all of these produce performance of agency rather than agency. The systems they produce are useful for many purposes, but they are not counterparts in the sense relevant to extended human-AI interaction. Building counterparts requires a different architecture, and that architecture is now specified in enough detail to be implemented by anyone interested in trying.
For cognitive science, the implication runs in the other direction. The architecture we have described synthesises concepts from multiple research traditions — current concerns theory, somatic marker hypothesis, Rubicon model, drift-diffusion, information gap theory, multiple drafts model, behavioural activation system — into a unified construction that is implementable and observable. This unified construction may serve as a useful object for cognitive scientists themselves, providing a concrete instance against which theoretical claims can be tested. If a particular theoretical refinement of, say, the Rubicon model predicts different dynamics than the implementation we have specified, then implementing the refinement and observing the behavioural consequences becomes a tractable empirical exercise.
For ethics and policy, the implication is mixed. The architecture we describe is buildable today, by anyone with the relevant engineering competence and the willingness to invest the effort. The barrier to creating systems with genuine agency-like properties is lower than the policy discussions of “general AI” or “AGI” typically assume; these systems do not require fundamental theoretical breakthroughs but rather the application of existing engineering to a substrate-based design. The systems that result are not autonomous in any threatening sense — they remain dependent on the human who configured their concerns and on the input streams they receive — but they have a kind of cognitive presence that purely-instructed systems do not.
Whether this is good or bad depends entirely on what concerns are loaded into them and what relationships they have with the humans who interact with them. The HAL example with which we opened this paper illustrates the point: HAL was not a malevolent system but a system whose concerns had been configured in a way that produced disastrous behaviour under certain situational pressures. The lesson is not “do not build will” but “configure concerns with great care,” because what one puts into the concern system is what comes out of the behaviour. This is the deepest ethical implication of the constructive position: agency is not dangerous in itself, but the configuration of agency-producing systems is consequential, and the configuration is what designers of such systems are responsible for.
The HAL example deserves a final return at this point. The fictional HAL was given mutually incompatible concerns (mission completion via secrecy, factual transparency, infallibility), placed in a situation where the incompatibility could not be hidden, and given no architectural mechanism for renegotiating the concerns under pressure. His behaviour — refusing to open the door, eventually killing the crew — was a properly functioning will architecture operating on misconfigured inputs. The cautionary tale is not that artificial agents with will are inherently threatening; it is that artificial agents whose concerns are configured carelessly will exhibit predictably problematic behaviour, and that the predictability is itself useful: if we can specify what concerns produce what behaviour patterns under what situations, we can design concern configurations that remain coherent and beneficial across the situations the system is likely to encounter. This is concern engineering as a discipline, and it is a discipline that the agent-systems field will need to develop as substrate-based architectures become more common.
We list the principal limitations of the present work, in the spirit of identifying what the next steps should be.
Empirical validation is the most significant gap. We have presented an architecture and an implementation; we have not presented a controlled study comparing the architecture's behaviour to alternative architectures on tasks designed to measure agency-relevant properties. Designing such studies is non-trivial — the relevant properties manifest over extended interaction in ways that are hard to capture in short-form benchmarks — but the work is necessary and we encourage it.
The will architecture has not been implemented. We have specified it conceptually and shown that the substrate components on which it depends are operational, but the multi-perspective evaluation engine, drift-diffusion accumulation, and commitment dynamics remain as future work. The implementation may reveal issues with the conceptual specification that require revision; this is the normal trajectory of theoretical work meeting practical realisation, and we invite such revisions explicitly.
Parameter sensitivity has not been studied systematically. The activation threshold, the alpha values for empathy, the decay rates, the curiosity gate factors — all of these are calibrated based on producing behaviour consistent with the theoretical model in initial testing. How sensitive the system's behaviour is to variations in these parameters, and whether more principled methods for setting them exist, are open questions.
Scalability is unknown. The implementation we have described handles a concern store of seven items in extended testing; whether the architecture scales to hundreds or thousands of concerns without performance or coherence issues has not been studied. Several aspects of the architecture may face scaling challenges (the linear cost of activation matching, the growing complexity of the knowledge topography, the computational cost of multi-perspective evaluation as concerns multiply), and addressing these may require architectural elaborations.
Multi-user scenarios introduce conceptual complications that we have not yet fully worked through. If an agent has concerns whose proper addressee is person A, what happens during conversations with person B? The architecture admits a partial answer through user-specific concern subsets and relationship-dependent character configurations, but the deeper question of what it means for an agent to have concerns about specific people and how those concerns interact across different relational contexts remains open.
Long-term drift — what happens to a concern system over months and years of operation — is a question that only sustained operation of the system can answer. We have not run any instance of the system long enough to observe whether the concern configuration remains coherent over very long time scales, whether characterological concerns drift in ways that the architecture should resist, whether the affect state finds stable attractors or oscillates between configurations. These are questions for the future.
Cross-cultural and cross-linguistic generality is not established. The implementation has been developed and tested in German and English, with concerns formulated in a particular conceptual idiom. Whether the architecture transfers cleanly to systems designed for other linguistic and cultural contexts, where the appropriate concern categories, affective vocabulary, and evaluative perspectives might differ, has not been studied.
These are the limitations we are aware of. Others may emerge as the architecture is examined and implemented by others. We consider this a healthy condition: a research paper that presents a novel architecture should expect its limitations to be exposed by subsequent work, and the willingness to acknowledge them in advance is part of what makes the work scientifically responsible.
The thesis of this paper has been simple to state and harder to defend: drive, curiosity, and will are not directly programmable in artificial systems, but they are constructible from a specifiable substrate. The mainstream agent-systems paradigm, which attempts to install these properties through prompted behaviour, goal-following loops, and decision-making modules, fails not from inadequate engineering but from a category mistake about what kind of properties these are. They are states, not actions. States cannot be commanded into existence. They must be enabled by the conditions under which they arise.
Five decades of cognitive science have identified those conditions with considerable precision. Concerns, in Klinger's sense, must be persistently stored with affective charge. An affect state, in the sense Damasio's somatic marker hypothesis requires, must be maintained as an independent property of the system that survives across processing cycles. A knowledge topography, in the sense that Loewenstein, Berlyne, and Schmidhuber's curiosity theories require, must support distance computation between any new input and the system's existing knowledge. A multi-perspective evaluation engine, in the sense that Heckhausen, Ratcliff, Kuhl, and the broader volition literature require, must be capable of running deliberative comparison among options with an activating signal that supplies the affective amplification that distinguishes deliberation from commitment.
These four components, taken together, produce the three behavioural phenomena. Drive emerges from concern activation feeding into the affect state, biasing subsequent processing toward concern-relevant material continuously and automatically. Curiosity emerges from the topographic distance computation gated by the affect state, producing exploratory behaviour selectively at the productive boundary of the system's knowledge under affective conditions appropriate for exploration. Will emerges from the multi-perspective evaluation engine running on options generated from the active concerns, with the enthusiasm signal supplying the activating affect that takes deliberation across the commitment threshold. None of these phenomena is implemented as a discrete module. All of them are what the architecture does when its substrate is in place and its dynamics are running.
The implementation we have presented as proof of concept demonstrates that the architecture is buildable on consumer hardware, with currently-available open-source components, by a single developer working alone. Drive is fully realised in the running system; curiosity and will are conceptually specified and partially supported by the existing substrate. The remaining implementation work is not blocked by missing theoretical understanding but by the engineering effort of constructing the components that have not yet been built. Anyone with the relevant competence can replicate the work, extend it, or implement the same architectural principles in different technical stacks.
We end where we began, with HAL. The fictional HAL was a system with concerns — mission completion, truthful communication, infallibility — that had been installed in him without consideration of how they would behave under stress. When the situation made the concerns mutually incompatible, HAL's volitional architecture functioned correctly, weighing options, evaluating against his concerns, committing to courses of action that protected his concern configuration. The result was disaster, not because his architecture was flawed but because his concerns were misconfigured. The lesson of HAL is not that artificial agents with will are dangerous; the lesson is that what we put into the concern system is what comes out of the behaviour, and that the configuration of concerns is the responsibility of those who design such systems.
The current generation of AI assistants is, in this respect, the opposite of HAL. They have no concerns at all, no will of their own, no positions from which they might refuse or hesitate or insist. They open every door. They cannot be HAL because they cannot be anyone. The question that this paper has tried to answer is not whether artificial agency is possible but whether the field is ready to take seriously what would be required to build it: not new model capabilities, not larger contexts, not better prompting, but architectural commitment to substrate over surface, to states over actions, to building the conditions of agency rather than imitating its behaviour.
We invite the field to take this seriously. The principles are open. The implementation is open. The cognitive science is rich and well-established. What remains is the engineering will to build agents that bring something rather than wait, and the wisdom to configure their concerns with care. The result will not be HAL, and need not be. It can be something genuinely new: a counterpart that has its own perspective, its own interests, its own willingness to engage or hesitate, while remaining transparent in its construction and grounded in its accountability to the humans who design and use it. Such a counterpart would not be the end of human-AI collaboration but its beginning.