Talk All Audiences 16:50 - 17:20 August 07, 2026

Meitar Ronen

Coding agents increasingly produce code that looks clean to static analysis, but still ship business-logic bugs: the kind threat modeling is supposed to prevent. Across 86 controlled experiments, we varied the security context given to gpt-5.5-codex across eight artifact families and three prompt structures, using the same codebase, PRD, and four planted business-logic flaws. Generic artifacts - DFDs, component diagrams, OWASP checklists - did not reliably help; several performed worse than no artifact at all. In one run, the agent - handed an OWASP checklist - cited A07 in its plan and shipped a 365-day bookmarkable magic link, the exact vulnerability A07 names. One pattern worked across every replication: feature-specific threats paired with concrete countermeasures, in a plan-forcing prompt. The talk walks through the mechanism behind these failure modes and the open-source tool we built to generate the artifact shape that worked.

Meitar Ronen

AI Research Lead at Clover Security

Meitar Ronen is AI Research Lead at Clover Security, where she drives the research direction for agentic development and AI infrastructure. She previously built computer vision algorithms and holds an M.Sc. in computer science with a focus on deep learning, with publications at CVPR and ICCV. She turned to the dark side and now leads AI research for cyber, focused on the hard problems that decide whether security agents can be trusted with real work.