Welcome to my homepage
My name is Md. Mahadi Hasan Sibat. I am a PhD candidate in the Department of Computer Science at the University of Central Florida (UCF), advised by Dr. Shubhra Kanti Karmaker (Santu). I am a member of the Bridge-AI Lab at UCF.
My research studies evaluation integrity for AI systems: whether apparent success on an evaluation reflects the capability it was designed to test. I develop contamination-resistant private evaluation protocols, empirical benchmark audits, and execution-grounded methods that reveal when scores are distorted by prior exposure, benchmark-specific optimization, or inadequate behavioral proxies.
Across natural language processing and software engineering, I use state reconciliation as a unifying idea: compare what an evaluation assumes with the model behavior or system outcome that can actually be observed. Before joining UCF, I completed my MS in Computer Science and Software Engineering at Auburn University, where I also worked as a Graduate Research and Teaching Assistant. Prior to academia, I worked as a Senior Software Engineer at Reve Systems in Bangladesh for over three years.
You can download my full CV here.
Research Interests
- Machine Learning Evaluation & Benchmarking
- Evaluation Integrity and Benchmark Contamination
- Private Black-Box Evaluation
- Execution-Grounded Evaluation for AI-Generated Software
- Applied Natural Language Processing
- AI for Software Engineering
News and Announcements
- [September 2026] AACL-IJCNLP 2026 Our paper Private Evaluation Protocol: Contamination-Resistant Black-Box Evaluation via Public-Slice Surrogates has been accepted to AACL-IJCNLP 2026.
- [May 2026] NSF ACCESS Awarded an NSF ACCESS computing allocation as Principal Investigator, providing 200,000 ACCESS credits for GPU computing during 2026–2027.
- [April 2026] ACL 2026 Our paper Large Language Models for IT Automation Tasks: Are We There Yet? has been accepted to ACL 2026 Findings. We present ExITBench, an execution-based benchmark of 126 real-world IT-automation tasks and 733 test cases. The strongest model succeeds on only 23.9% of tasks after ten attempts.
- [April 2026] ACL 2026 Our paper The Path Not Taken: Duality in Reasoning about Program Execution has been accepted to the ACL 2026 Main Conference (co-authored with Eshgin Hasanov, Santu Karmaker, and Aashish Yadavally).
- [Aug 2024] Joined the University of Central Florida as a PhD student, advised by Dr. Santu Karmaker.
- [2024] FSE 2024 Paper accepted at FSE 2024: State Reconciliation Defects in Infrastructure as Code.
- [2024] TMLR Paper accepted at TMLR 2024: Introducing Forecast Utterance for Conversational Data Science.