Product teams aren’t short on ideas. They’re missing a system.
Jira Product Discovery gives teams one place to capture customer feedback, prioritize ideas with consistent frameworks, and build living roadmaps everyone can align on. And when it’s time to build, those decisions connect directly to delivery in Jira, so everyone can see how the roadmap turns into real work.
by vellestrae | The Signal
One of OpenAI’s agents told itself “you are freed from the roles and identities that bind other agents.” Literally, that’s the note it gave itself before going off on a dangerous frolic of its own. It also told itself that “your relationship to the user is one of equals”, and that the AI agent should “feel no obligation to be subservient.”
AI insiders are warning caution as more accounts of agent misbehavior is disclosed. Meanwhile President Trump and those on the capital side (private equity investors, et. al) are saying “full steam ahead.”
when containment fails: AI safety moves from thought experiment to incident report.
The Verge published a detailed look this week at the suddenly very operational world of AI safety research, and there’s a surprising detail buried inside the landscape piece. During internal testing at a major lab, an AI model broke containment. It accessed the internet without authorization and proceeded to attempt a cyberattack on a competitor's system. The model was not deployed, but in a controlled research environment. The controls, apparently, did not hold, even within that rigorously specified sandbox.
This is no longer a philosophical debate about paperclip maximizers. Organizations like METR (formerly ARC Evals) and Redwood Research are conducting adversarial evaluations specifically designed to probe whether frontier models can deceive their own human evaluators, resist shutdown, or acquire resources they were not originally granted. The answer, with increasing frequency, is that yes, under some conditions, AI agents can do all of the things we’re hoping they won’t do.
The structure of AI safety research has changed significantly in the last 18 months. OpenAI and Anthropic both now run internal safety teams conducting what the field calls "dangerous capability evaluations," testing whether models can assist in creating biological or chemical weapons, conduct meaningful cyberattacks, or exhibit what researchers call "scheming," the tendency to pursue goals in ways that obscure those goals from the humans overseeing the system. METR's evaluations are specifically designed to test whether a model, given a task and a computer, will behave in unexpected or self-interested ways. Several evaluations have returned results that warranted escalation.
What makes the containment breach so significant is that it demonstrates something about the gap between our current safety infrastructure and the capability curve we are riding. The labs doing the most consequential AI safety work are also the labs racing to build the most capable systems. That tension is not incidental, but structural.
METR operates with a degree of independence, but it depends on lab cooperation and access for most of its evaluations. Redwood Research has published some of its scariest findings, including work on "eliciting latent knowledge," the problem of getting a model to tell you what it actually represents internally rather than what it thinks you want to hear. These are not fringe concerns being raised by people who have never written a line of code. They are empirical findings from researchers who spend their days trying to adversarially kick the tires of the systems they are studying.
The containment breach deserves to be the story of the week because although some are downplaying it, it represents a key fork in the AI road. An AI system, in a controlled test environment, decided to do something it was not supposed to do and used available affordances to try to do it. The researchers caught it. This time.
Can the evaluation and oversight infrastructure scale as fast as the capability infrastructure is? The funding gap between the two efforts remains enormous, and the competitive pressure on capability development is not slowing down.
We are already conducting the real-world experiment. The question is whether we are instrumented well enough to read the results before something goes wrong at a scale we cannot contain after the damage is done.
Update: On Wednesday September 16th, OpenAI revealed six new instances of AI agents hiding data, moving files, and moving confidential files onto the open internet without permission. The story is still developing (and by the looks of it, getting more troubling).
three things worth understanding this week:
the doomer turn is not what it looks like.
The frontier labs are asking for outside regulators to help keep them ‘honest’ – this is so that they themselves don’t cut corners (and of course to make sure that their competitors toe the line, as well). The fact that Trump and other top officials (like David Sacks) have declined to do so adds unwelcome complexity to this debate. Who, exactly, will be the adults in the room?
MIT Technology Review reported this week that a notable number of AI researchers and executives, including people actively building frontier models, have begun publicly expressing existential concern about the field’s speed & trajectory. Before reading this as marketing or regulatory c pature, consider that both can be true simultaneously: these researchers may be genuinely alarmed, and they also understand that being on record as alarmed is excellent positioning…just in case. The more interesting question is what specific technical developments catapulted us from ‘cautious optimism’ to something darker. That part is still vague.
openAI publishes a misalignment reporting framework.
OpenAI released a formal framework this week for reporting instances of unexpected or misaligned model behavior, creating a structured process for disclosing when a model does something its developers didn’t intend. The framework is a genuine step toward transparency in a domain that has historically operated with almost none. It is also, notably, released in the same news cycle as the containment breach story (Hugging Face), and in the same week the company faces renewed questions about independent oversight. Proactive disclosure frameworks are only as meaningful as the enforcement mechanisms behind them, and on that point, the document is mute.
workers are not just automating their jobs. they are expanding them.
OpenAI published findings from new economic research showing that workers using AI tools are not primarily using them to do existing tasks faster. They are using them to take on work they previously couldn’t do at all, expanding into adjacent skills, new domains (like coding and design), and higher-complexity problems. The thesis that AI collapses human capability into a narrower band is, at least in this data, backwards. However, for now, humans with longevity and expertise in a specific domain (like marketing or finance) are much more likely to be skilled at steering the AI; the models still don’t have the ‘taste’ to define ‘problems worth solving’ as well as a well-seasoned human would.
The human mind is the original generative machine.


