Welcome to my homepage

My name is Md. Mahadi Hasan Sibat. I am a PhD candidate in the Department of Computer Science at the University of Central Florida (UCF), advised by Dr. Shubhra Kanti Karmaker (Santu). I am a member of the Bridge-AI Lab at UCF.

My research studies evaluation integrity for AI systems: whether apparent success on an evaluation reflects the capability it was designed to test. I develop contamination-resistant private evaluation protocols, empirical benchmark audits, and execution-grounded methods that reveal when scores are distorted by prior exposure, benchmark-specific optimization, or inadequate behavioral proxies.

Across natural language processing and software engineering, I use state reconciliation as a unifying idea: compare what an evaluation assumes with the model behavior or system outcome that can actually be observed. Before joining UCF, I completed my MS in Computer Science and Software Engineering at Auburn University, where I also worked as a Graduate Research and Teaching Assistant. Prior to academia, I worked as a Senior Software Engineer at Reve Systems in Bangladesh for over three years.

You can download my full CV here.

Research Interests

  • Machine Learning Evaluation & Benchmarking
  • Evaluation Integrity and Benchmark Contamination
  • Private Black-Box Evaluation
  • Execution-Grounded Evaluation for AI-Generated Software
  • Applied Natural Language Processing
  • AI for Software Engineering

News and Announcements

  • [September 2026] AACL-IJCNLP 2026 Our paper Private Evaluation Protocol: Contamination-Resistant Black-Box Evaluation via Public-Slice Surrogates has been accepted to AACL-IJCNLP 2026.
  • [May 2026] NSF ACCESS Awarded an NSF ACCESS computing allocation as Principal Investigator, providing 200,000 ACCESS credits for GPU computing during 2026–2027.
  • [April 2026] ACL 2026 Our paper Large Language Models for IT Automation Tasks: Are We There Yet? has been accepted to ACL 2026 Findings. We present ExITBench, an execution-based benchmark of 126 real-world IT-automation tasks and 733 test cases. The strongest model succeeds on only 23.9% of tasks after ten attempts.
  • [April 2026] ACL 2026 Our paper The Path Not Taken: Duality in Reasoning about Program Execution has been accepted to the ACL 2026 Main Conference (co-authored with Eshgin Hasanov, Santu Karmaker, and Aashish Yadavally).
  • [Aug 2024] Joined the University of Central Florida as a PhD student, advised by Dr. Santu Karmaker.
  • [2024] FSE 2024 Paper accepted at FSE 2024: State Reconciliation Defects in Infrastructure as Code.
  • [2024] TMLR Paper accepted at TMLR 2024: Introducing Forecast Utterance for Conversational Data Science.