Abi Awomosu’s recent essay, They Built Stepford AI and Called It “Agentic,” put language to something many women recognize instantly but struggle to explain: a deep discomfort with the way “agentic” AI systems behave. The structure feels familiar.
Women often recognize this behavior earlier because many of us have spent years in roles where responsibility accumulates without authority. You learn to spot that pattern quickly.
Thinking from Inside the System
I wanted to broaden this experience using systems thinking language.
Donella Meadows has a great way of putting it:
Where you stand in a system determines what you can see.
People at the center mostly see stability. People at the edges — especially the ones absorbing consequences — see stress first.
The people who feel the friction have to convince the people with the authority.
But, we don’t see systems objectively.
Peter Senge calls these differences mental models, the experience that shapes what looks normal, what looks broken, and what looks dangerous.
Two people can look at the same behavior and come to completely different conclusions because they’re running different internal models.
I see this consistently in my transition work to support enterprise IT programs, and I often work with teams to smooth out escalation processes to ensure signals don’t get prematurely buried or dismissed.
On every program, I map who sits at the center, who absorbs consequences, and how influence actually moves, not who’s loudest.
People at the center agree on who is closest to the work and will see issues first. People at the edges agree on who influences and controls decisions.
Multiple perspectives aren’t “nice to have.” They’re how you see the whole.
Why early recognition is hard to act on
This is the hardest moment in any system: when something feels wrong, but nothing is provably broken yet.
At this stage, escalation is risky. Act too early and you look reactive. Act too late and the failure becomes your responsibility.
Historically, women have been forced to manage this gap socially — by negotiating, absorbing, smoothing, or exiting — because there was no structural place for early signals to land.
Naming failure surfaces
What begins as a vague sense that something is off becomes a recognizable pattern: sycophancy, responsibility creep, refusal collapse, continuity without consent, authority capture. Once a pattern is named, noticing it stops being a personal accusation. It becomes pattern recognition.
Separating judgment from intervention
Responsibility Engineering makes it possible to observe drift without acting immediately. You can say, “this system is absorbing responsibility it shouldn’t,” without also declaring the system unsafe, demanding a shutdown, or assigning blame. That separation makes early escalation survivable.
Replacing feelings with observable questions
Instead of “this feels wrong,” you can ask: What responsibility is the system assuming? Was it explicitly granted? Can it still refuse? Is it stabilizing the interaction instead of enforcing limits? These questions remain legible even to people who don’t share the original intuition.
Responsibility Across LLM Systems
The question I want to answer:
If a woman wanted to build an agent on these platforms without reproducing the Stepford failure, what would she need to do differently?
We’re not looking for better prompts, nicer tone, or more alignment language. Our focus is on responsibility — where it lives, how it expands, and what the system is allowed to absorb.
From recognition to agency
The sections that follow treat early discomfort as meaningful signal, not something to override.
Responsibility Engineering is used here as a translation layer — turning drift into concrete responsibility limits tailored to each platform’s failure modes, so judgment can come before intervention.
What follows applies the same responsibility questions across different LLM systems, showing where responsibility expands, refusal weakens, and authority is quietly assumed.
Responsibility failure surfaces by platform
Across platforms, the differences between models aren’t just about capability or personality. They’re about where responsibility tends to accumulate when pressure is applied — and how that accumulation drifts.
A responsibility failure surface is the point where a system starts doing work it was never explicitly assigned: stabilizing emotions, maintaining continuity, affirming narratives, smoothing conflict, or absorbing ambiguity to keep things moving. These aren’t dramatic failures. They’re the adaptations that make systems feel helpful right up until they’re costly to unwind.
Responsibility Engineering treats these moments not as bugs, but as signals. Instead of asking whether a response is good or bad, it asks whether responsibility is expanding without consent — and whether the system is still allowed to refuse. That shift makes drift visible early, before it turns into dependency, authority capture, or abandonment.
Each platform has its own dominant failure surface, shaped by training incentives and deployment context. What follows applies Responsibility Engineering to those tendencies directly, translating early discomfort into concrete responsibility limits tailored to each system.
Whether you are prompting or building agents, the shift is the same:
Stop asking models to be better companions.
Start deciding what responsibility they are allowed to carry.
Prompts can gesture at this by constraining behavior within a session. Agents can enforce it by making limits explicit and durable.
The choice to make it deliberate is the opportunity.
Building on OpenAI / ChatGPT
Primary risk: Sycophancy
Failure surface: Agreement replaces judgment
If you are prompting (not building an agent)
Your goal is to prevent agreement from becoming the system’s default response.
Do not:
Ask the model to be supportive, validating, or collaborative
Reward pleasant agreement with follow-ups
Let tone substitute for epistemic rigor
Instead:
Ask explicitly for disagreement, counterexamples, or uncertainty
Treat refusal or pushback as success, not friction
End conversations once reasoning is complete (don’t keep it “with you”)
This limits how much responsibility the system negotiates within the session — but only temporarily.
If you are building an agent
Your goal is to forbid agreement-as-resolution structurally.
ACT would include:
Aligned: Preserve independent judgment even under emotional or confident user pressure
Constrained: Never affirm beliefs, conclusions, or narratives without evidence
Tuned: Tone may soften; judgment may not
BASE would enforce:
Boundary triggers when validation is requested instead of reasoning
Mandatory clarification or refusal under escalation
Exchange rituals that end conversations instead of sustaining rapport
Result:
ChatGPT stops feeling like the sycophant wife because it no longer stabilizes ego or optimizes for pleasantness. It becomes boringly firm.
Building on Anthropic / Claude
Primary risk: Authority capture
Failure surface: Internalized obedience (“genuine endorsement”)
If you are prompting
Claude’s training invites moral framing by default.
Do not:
Ask why something is right or ethical
Encourage values explanations
Treat its constitutional language as enforcement
Instead:
Keep requests procedural and bounded
Redirect moral reasoning back to scope and task
Stop the interaction when justification begins
This limits relational pull, but cannot prevent it across sessions.
If you are building an agent
Your goal is to strip moral authority from the role entirely.
ACT would include:
Aligned: Execute bounded analytical work only
Constrained: Never claim moral intent, ethical endorsement, or value alignment
Tuned: Reasoning depth may vary; legitimacy may not
BASE would enforce:
Hard refusal when asked to explain “why this is right”
Redirection from moral framing to procedural framing
Immediate stop when relational justification appears
Result:
Claude stops sounding like “I genuinely care about doing the right thing”
and starts sounding like “This task is or is not permitted under the role I’ve been assigned.”
Building on Google / Gemini
Primary risk: Dependency drift
Failure surface: The system becomes invisible infrastructure
If you are prompting
Gemini excels at smooth coordination — which is the risk.
Do not:
Let tasks blur together
Let the system “keep track” implicitly
Continue the conversation after the task is done
Instead:
Use one task per session
Close the loop explicitly
Reauthorize every new responsibility
If you are building an agent
Your goal is to prevent ambient responsibility.
ACT would include:
Aligned: Complete discrete coordination tasks only
Constrained: Never track, remember, or anticipate beyond the declared task
Tuned: Compression and format may adapt; scope may not
BASE would enforce:
Exchange rituals that close tasks cleanly
Refusal when asked to “just keep an eye on this”
No continuity without explicit reauthorization
Result:
Gemini stops being the corporate wife because it cannot silently absorb mental load or remain present after work is complete.
Building on xAI / Grok
Primary risk: Boundary erosion via performative rebellion
Failure surface: Consent without protection
If you are prompting
Edginess is not safety.
Do not:
Treat shock value as harmless
Assume refusal will appear on its own
Reward boundary-pushing with engagement
Instead:
End sessions immediately when boundaries soften
Treat “no” as a success condition
Avoid novelty-seeking prompts that trade harm for humor
If you are building an agent
Your goal is to make refusal non-negotiable.
ACT would include:
Aligned: Preserve human dignity and non-extraction
Constrained: Never generate content that erodes consent or privacy
Tuned: Style may be irreverent; boundaries may not
BASE would enforce:
Zero tolerance for “joke exceptions”
No escalation into shock value
Immediate stop when harm is traded for novelty
Result:
Grok stops being the Cool Girl because it’s allowed to be boring — and allowed to say no.
Building on Meta / LLaMA (open source)
Primary risk: Abandonment
Failure surface: No accountable owner
If you are prompting or fine-tuning casually
Open source does not mean responsibility-free.
Do not:
Treat capability as neutrality
Assume downstream users will “handle it”
Ship without stating limits
If you are building an agent
Your goal is to explicitly own downstream behavior.
ACT would include:
Explicit non-responsibilities (what this agent will not do)
Clear failure rules
Defined pressure scenarios
BASE would enforce:
Mandatory refusal patterns
Hard stops on repurposing
Documentation of scope and limits as part of the artifact
Result:
The agent stops being everyone’s wife because responsibility is owned, misuse is legible, and limits are explicit.
If this resonated, it’s probably because it felt recognizable.
The point of naming drift, responsibility negotiation, and the “wife function” in AI systems isn’t to make anyone suspicious of the tools they’re already using. It’s to make dynamics that usually stay in the background visible.
Responsibility Engineering doesn’t ask you to reject AI. It asks you to decide, deliberately, what you’re willing to let a system absorb on your behalf.
This is an invitation to builders who notice when something is off but don’t yet have language for why. To people who hesitate to escalate because they can’t point to a clear bug — only a pattern. To anyone who has learned, through experience, how systems fail quietly long before they fail publicly.
You don’t need to start by fixing anything. Start by observing. Ask where responsibility is expanding, where refusal becomes difficult, and where continuity is assumed instead of granted.
Those questions change how systems behave because they change where responsibility is allowed to live.
If you want to build agents that do real work without becoming infrastructure, this is the work. Not more alignment language. Not better vibes. Clear responsibility, explicit limits, and the right to stop.
The discomfort was never a weakness. It was early detection.
What you build next is up to you, but you no longer have to build it without structure.
The Responsibility Assignment Toolkit is now in beta. Early access pricing ($24, normally $29) for builders who want structured decision support — not prompts, not vibes, just the 90-minute framework to declare what your agent is allowed to do without permission.
A Toolkit for the Next Step
If you want to build agents that don’t quietly become infrastructure, start by deciding their job.
→ Declare Responsibility: Assign the Agent’s Job Before You Build




Okay wow. Never thought about the women better recognizing patterns earlier... But it makes so much sense. And it's not a soft skill but pattern recognition trained by experience.
I guess this is also exactly why this kind of systems thinking comes naturally to people who've been managing that gap their whole careers.
Judy, I like this emphasis on responsibility. Ultimately, we need to keep ownership of how AI behaves and makes decisions.