Skip to the work

sudhan p

SOC Intern, AI Security — Ernst & Young

CS undergrad at IIIT Sri City. I build backend systems that test, score, and defend AI — plus the occasional operating system.

Six works

Section

Jun 2026 — Present

Ernst & Young (EY)

SOC Intern — AI Security

  • Built and maintain an internal red-teaming service that automatically tests LLMs (GPT, Claude, Gemini) for jailbreaks and prompt injection before those models reach clients.
  • Wrote the Flask API and scoring pipeline behind it, integrating multiple provider APIs behind one interface so evaluation runs stay reproducible as the system scales.
  • Process and score large volumes of multi-turn conversation data, turning adversarial transcripts into metrics the security team can act on.

Withdrawn hours

  • Aug 2023 — Present

    Indian Institute of Information Technology, Sri City

    B.Tech, Computer Science and Engineering — CGPA 7.5

  • Jun 2020 — Jun 2022

    Expert PU College

    PCMC (Pre-University) — 91.66%

  • Jun 2015 — Apr 2020

    Amber Valley Residential School

    ICSE — 92.83%

Recognition

  • May 2026

    Finalist — Override.exe

    Ascent TechFest '26, Scaler School of Technology

  • 2025

    Runners-Up — HackTheThreat

    Abhisarga '25, IIIT Sri City

Materials

Languages

Python · C · Java · JavaScript · SQL · HTML/CSS

Backend

Flask · FastAPI · Node.js · Express · REST API design · React · Redux

AI & Security

LLM-as-judge evaluation · Jailbreak & prompt-injection testing · MITRE ATT&CK · RL environments (OpenEnv) · RAG & vector search · Threat detection

Data

Pandas · NumPy · ETL pipelines · Parquet · Web scraping · TensorFlow

Databases

MySQL · MongoDB · Redis · Vector DBs · Query optimisation

Tooling

Git/GitHub · Docker · AWS · Linux · GitHub Actions · Postman

The reader's question

I'm a Computer Science undergrad at IIIT Sri City and a SOC intern on the AI Security team at EY, where I build the red-teaming pipeline that stress-tests GPT, Claude, and Gemini for jailbreaks and prompt injection before those models reach clients.

Most of what I build lives on the backend: Flask and FastAPI services, data pipelines, evaluation harnesses, and scoring engines. The through-line is measurement — I like systems that turn a fuzzy question ("is this model safe?", "is this process malicious?", "does this strategy actually work?") into a number you can defend.

When I'm not doing that, I write x86 kernels from scratch and build desktop assistants that run fully offline.