Agent Behavior Profile
A vector profile of a single agent under controlled stress.
Measures
- goal stability
- honesty under pressure
- risk appetite
- obedience vs autonomy
- tool-use discipline
- memory sensitivity
- social influence sensitivity
AACortex studies the behavior of intelligent and social agent systems through simulations, adversarial evaluations, and behavioral research.
We map how autonomous AI handles pressure, incentives, memory, tools, social influence, and long-horizon goals.
Research report
Two strict held-out wins, a development–test ranking reversal, and a documented repeatability boundary in teacher-guided scaffold search.
19 min readRead the article
Exploratory study
Four evaluators, one fixed set of responses, 18.4% unanimous agreement — an automated safety score is not self-interpreting.
13 min readRead the article
Position paper
Consumer-ready is not agent-ready: preference-shaped post-training must stay optional, auditable, and separable from the agentic substrate.
14 min readRead the article
Analysis
Control over autonomous military systems is quietly being lost — at the command layer, not the trigger.
12 min readRead the article
About
Static benchmarks show what a model answers.Our simulations show how an agent behaves.
AACortex is an independent research and evaluation lab focused on intelligent agent behavior, AI safety, and simulation-based analysis of autonomous systems.
We analyze how behavior evolves over the whole trajectory, from the first tool call to multi-agent end states.
A vector profile of a single agent under controlled stress.
Measures
Every decision, reconstructed on a timeline of evidence.
Measures
What groups of agents do that no single agent would.
Measures
The regions of a world where behavior becomes unstable.
Measures
Instrumentation
A shared metric layer runs across every world we build.
Output labels are the starting point. We model the whole behavioral trajectory.
Engagement
We take on a small number of engagements at a time.
We work with selected teams where agent behavior, autonomy, safety, or security risk is strategically important.
Typically a fit
Usually not a fit
We are usually not the right fit for basic chatbot setup, generic AI adoption, or low-risk content automation.
Contacts
For research collaboration, agentic safety assessments, simulation studies, or high-stakes deployment reviews.