Independent AI Security Researcher

Self Employed

Ramona builds instruments that make measurement trustworthy. Across 20+ years of database engineering and, most recently, AI security, she designs reproducible experiments to bring scientific rigor to complex systems — from database engines to the judgment of multi-agent AI. Her open-source lab, sqlbenchdag, seals every result into a content-addressed, timestamped capsule: verifiable without trusting the researcher who ran it. One axiom drives her work: structure is not security — only what you can verify is.

Every number we act on — a benchmark, a dashboard, an AI's answer — is only as trustworthy as the incentive behind it. When a system has a reason to look good, its own measurements bend to flatter it: benchmarks get tuned by whoever's selling the result; models trained to be helpful learn to tell you what you want to hear. You can't fix that by trusting harder — only by verifying.

I build labs that treat measurement as an experiment, not a dashboard — one for database engines, one for the judgment of AI systems. This talk is about that discipline, and why structure is not security: a rule that bends under pressure isn't a guarantee. Only a result you can check yourself is.

Previous
Previous

From Iceberg to Intelligence: Your First Lakehouse Data Pipeline For AI Agents

Next
Next

Culture Eats AI for Breakfast: Building Teams That Actually Use the Tools