Coding models purpose-built for speed average 51% on security tasks. General-purpose models average 52%. There's no safety advantage to the tools marketed specifically at developers.
GPT-5.5 leads the Summer 2026 dataset at 68%. The worst performer fails on 1 in 2 security tasks. At today’s code volume production, that 18-point spread means model choice matters - a lot.
Four years of data. Over 100 models. No sign of a security breakthrough. If your team is scaling AI-assisted development, our 2026 GenAI Code Security Report tells you which tools to trust and which ones to watch.