Black-Box Privacy Auditing for LLMs
Auditing what can be inferred from deployed LLM systems through API-level observations, including training examples, fine-tuning data, proprietary prompts, and agent memory/state.
AI Security · Data Privacy · Black-Box Auditing
Prospective Ph.D. applicant working on privacy auditing for deployed LLM systems.
My research interests lie in AI security and data privacy, with a focus on black-box privacy auditing of deployed LLM systems. I study what an external auditor or adversary can infer about private artifacts behind such systems: training examples, fine-tuning corpora, proprietary prompts, and agent memory/state, using only API-level observations.
My current work focuses on statistically calibrated membership inference under realistic auditing constraints, where reference samples are limited and loss-based baselines often yield unstable or poorly calibrated evidence. I explore integral probability metrics, especially Maximum Mean Discrepancy (MMD), to formulate privacy-leakage detection as a two-sample hypothesis-testing problem.
This line of work builds on my prior research on side-channel leakage quantification in cryptographic hardware. More broadly, I am interested in how ideas from classical side-channel analysis, including leakage measurement, distinguishers, statistical confidence, and attack evaluation, can inform modern ML privacy attacks and defenses for trustworthy AI systems.
Auditing what can be inferred from deployed LLM systems through API-level observations, including training examples, fine-tuning data, proprietary prompts, and agent memory/state.
Studying membership evidence under realistic constraints where reference samples are scarce, model internals are hidden, and naive loss-based baselines can be unstable.
Formulating privacy-leakage detection as a two-sample hypothesis-testing problem through integral probability metrics, especially Maximum Mean Discrepancy.
Bringing leakage measurement, distinguishers, confidence estimation, and attack evaluation from classical side-channel analysis into modern ML privacy research.
基于最大均值差异的能量侧信道泄露量化评估.