Machine Learning & AI · Software Engineering
I'm a senior (B.Tech CSE, minor in Data Science) with real world experience in medical imaging, multi-omics cancer modelling and LLM alignment, plus ML infrastructure and systems work in Rust, C++ and Go.
Website · Resume · LinkedIn · Google Scholar · X
- Summer Research Intern(onsite) at Indian Institute of Technology (IIT) Kharagpur under Dr. Sourangshu Bhattacharya. May 2026 - Present.
- Research Intern(virtual) at Jadavpur University CMATER Lab under Prof. Debotosh Bhattacharjee. November 2025 - May 2026.
- Summer Research Intern(onsite) at New Jersey Institute of Technology, USA, in the Learning-based Decision Making Lab under Dr. Arnob Ghosh. June 2025 - November 2025.
- Zed Code Editor (9 merged PRs)
- Kubescape (8 merged PRs)
- OpenCV (5 merged PRs)
- Rust Language (2 merged PRs)
- Zen Browser (2 merged PRs)
- Activity Watch (2 merged PRs)
- Python (1 merged PR)
- Ghostty (1 merged PR)
- Root by CERN (1 merged PR)
- hnn-core by jonescompneurolab (1 merged PR)
- jenkins (1 merged PR)
- skopeo by podman-container-tools (1 merged PR)
- Kimi K3 for All : Built a from-scratch training pipeline for Kimi K3, a 2.8T-parameter Mixture-of-Experts model released with inference-only code, by fixing 4 undocumented bugs blocking gradient flow through its router and experts. Then trained a 1.27B-parameter (0.364B active) model with the pipeline. Weights on Hugging Face, write-up on my blog.
- Jupyter Extension for Zed Editor : A Zed extension that adds Jupyter notebook support. Zed extensions can't draw custom UI, so it works through what Zed already has: a language server that checks .ipynb files, and code actions to pair a notebook with a runnable # %% script, run all cells, or clear outputs.
- DebateBench : A multi-LLM debate platform. Anonymised agents argue a topic over several rounds, and each round's winning response becomes shared context for the next. Every agent gets a fresh random code each round, so the orchestrator judging them never learns which model wrote what and can't build a bias toward one.
- Amazon ML Challenge 2026 in a Self-Imposed 24 Hours: 0.988216 F0.5 : Matching 10 million noisy business records to 1.7 million businesses across US, India and an unseen France. September 2026.
- Kimi K3 for All : Moonshot never released a training pipeline for Kimi K3, only inference code. Rebuilding one from scratch, and the small model it produced along the way. August 2026.
- Recursive and Wrapper-Based Feature Selection for Breast Cancer Diagnosis and Prognosis. Ayushi Bhattacharjee, Arnesh Banerjee, Arpita Talukdar. 4th Analytics Global Conference (AGC 2026), March 2026. SPRINGER ASCAR Series. Oral presentation.
- Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment. Kartik Pandit, Sourav Ganguly, Arnesh Banerjee, Shaahin Angizi, Arnob Ghosh. 2025. arXiv: https://arxiv.org/abs/2510.03520
- Failure Modes of Large Language Models on Research-Level Mathematics: A Taxonomy and an Empirical Characterisation. Arnesh Banerjee, Ayushi Bhattacharjee. arXiv: https://arxiv.org/abs/2606.24902
- An Intelligent Weakly Supervised Framework for Breast Thermography Segmentation Using Hybrid CNN–Transformer Networks. Arnesh Banerjee, Debotosh Bhattacharjee. In preparation for Expert Systems with Applications.
Contact me: arneshbanerjee24 [at] gmail [dot] com




