tarun koyalwar

Making agents better at hacking.

AI researcher @ ProjectDiscovery. Benchmarks, evals, and observability that show what language models actually do on offensive-security tasks.

now

Right now: benchmarking and experimenting with OSS models, all things agents and models, at ProjectDiscovery.

First talk that will be recorded: BSides Las Vegas 2026, this August.

writing

see all

talks

Watching Agents Work: A Behavioral Audit of 189 Offensive-Security LLM Runs

bsides las vegas · ground truth · aug 2026 · will be recorded

From Mapping to Mitigation

black hat asia arsenal · apr 2025

see all

selected work

Neo: evals & benchmarking (external)

ProjectDiscovery's offensive-security AI agent, and the open question of whether an autonomous agent can actually hack. Builds the harness that measures it: evals at scale, benchmarking, and the trace observability that shows what the agent actually did on a target.

neo · evals · benchmarking · observability

Nuclei (external)

The vulnerability scanner much of offensive security runs on. Core team through the v3 era, now on Neo.

go · ~29k★ · core team, v3 era

Alterx (external)

Subdomain permutation generator driven by patterns instead of a static wordlist: define the patterns, get candidate hostnames to enumerate before a scan.

go · ~940★ · author

Talosplus (external)

Recon-automation framework in Go: plain bash scripts become a managed parallel execution graph. One of the two tools that landed the ProjectDiscovery job.

go · ~92★

see all

say hi

In Pune now, Las Vegas next month. X is the fastest line, faster than email. Always up for talking AI, agents, or travel.

say hi on X (external)