|
| 1 | +--- |
| 2 | +layout: default |
| 3 | +title: Muskaan Chopra |
| 4 | +description: false |
| 5 | +--- |
| 6 | + |
| 7 | +{: style="float: right; margin: 0 0 1em 1em; max-width: 250px; border-radius: 4px;"} |
| 8 | + |
| 9 | +I am a PhD researcher at the **Applied Machine Learning Lab (AML Lab)** at the **University of Bonn** and the **Lamarr Institute for Machine Learning and Artificial Intelligence**. |
| 10 | + |
| 11 | +My research is centered around a question that sounds simple, but turns out to be surprisingly difficult: |
| 12 | + |
| 13 | +**When should a machine learning model trust its own prediction — and when should it know that it does not have enough information?** |
| 14 | + |
| 15 | +I am particularly interested in **reliable and resource-efficient language models**, with a focus on small and compact models that can reason, recognize uncertainty, and make better decisions under limited information or computational resources. My current work studies **context sufficiency, abstention, calibration, and selective prediction**, including how these behaviours emerge during training and whether the signals we can observe inside a model are actually used when it makes a decision. |
| 16 | + |
| 17 | +A second thread of my work looks at what happens when models become smaller or more efficient. I have worked extensively on **quantization and compact language models for critical error detection in machine translation**, studying where compression is essentially free and where it begins to affect reliability. More broadly, I am interested in evaluation settings where aggregate accuracy alone is not enough and individual mistakes can have very different consequences. |
| 18 | + |
| 19 | +Before moving towards language models, much of my research focused on **self-supervised learning and medical imaging**, particularly diabetic retinopathy screening. This continues to shape how I think about trustworthy AI: models should not only perform well, but should also communicate when their predictions are unreliable. |
| 20 | + |
| 21 | +Across these areas, I am especially interested in models that are **small enough to study carefully, efficient enough to deploy, and reliable enough to know their limits**. |
| 22 | + |
| 23 | +## Research interests |
| 24 | + |
| 25 | +- **Reliable & Trustworthy Machine Learning:** Understanding when models fail, when they should abstain, and how reliability can be evaluated beyond average accuracy. |
| 26 | +- **Small & Efficient Language Models:** Compact models, quantization, compression, and the relationship between model efficiency and behavioural robustness. |
| 27 | +- **Context Sufficiency & Abstention:** Studying whether language models can recognize when the available information is sufficient to answer, and how this signal influences their decisions. |
| 28 | +- **Mechanistic & Developmental Analysis:** Investigating where reliability-related signals are represented inside neural networks, whether they are causally used, and how they emerge during training. |
| 29 | +- **Selective Prediction & Calibration:** Designing systems that can defer uncertain or risky predictions instead of treating every input as equally answerable. |
| 30 | +- **Machine Learning for High-Stakes Applications:** Reliable evaluation in areas such as machine translation and medical AI, where seemingly small errors can have disproportionate consequences. |
| 31 | + |
| 32 | +## Selected recent work |
| 33 | + |
| 34 | +### Knowing When Not to Predict |
| 35 | + |
| 36 | +**Self-Supervised Learning and Abstention for Safer Diabetic Retinopathy Screening** |
| 37 | + |
| 38 | +*IJCAI-ECAI 2026* |
| 39 | + |
| 40 | +We study how self-supervised pretraining influences not only classification performance but also a model's ability to identify cases on which it should abstain. The work explores selective prediction as a way of moving beyond accuracy towards safer medical AI. |
| 41 | + |
| 42 | +### Towards Reliable Machine Translation |
| 43 | + |
| 44 | +**Scaling LLMs for Critical Error Detection and Safety** |
| 45 | + |
| 46 | +*ECIR 2026* |
| 47 | + |
| 48 | +We investigate how language models of different scales perform at detecting meaning-critical translation errors and examine the trade-offs between model size, reliability, and computational cost. |
| 49 | + |
| 50 | +[\[Paper\]](https://arxiv.org/abs/2602.11444) |
| 51 | + |
| 52 | +### How Small Can You Go? |
| 53 | + |
| 54 | +**Compact Language Models for On-Device Critical Error Detection in Machine Translation** |
| 55 | + |
| 56 | +*IEEE BigData 2025* |
| 57 | + |
| 58 | +This work explores how far language models can be compressed while retaining their ability to detect critical translation errors, with particular attention to parameter-efficient and quantized models. |
| 59 | + |
| 60 | +[\[Paper\]](https://arxiv.org/abs/2511.09748) |
| 61 | + |
| 62 | +### SynCED-EnDe 2025 |
| 63 | + |
| 64 | +**A Synthetic and Curated English-German Dataset for Critical Error Detection in Machine Translation** |
| 65 | + |
| 66 | +*ECIR 2026* |
| 67 | + |
| 68 | +We introduce a structured benchmark for critical error detection containing fine-grained error categories designed to support more systematic evaluation of both compact and large language models. |
| 69 | + |
| 70 | +[\[Paper\]](https://arxiv.org/abs/2510.05144) |
| 71 | + |
| 72 | +### Functional Knowledge Transfer with Self-Supervised Representation Learning |
| 73 | + |
| 74 | +*IEEE International Conference on Image Processing (ICIP), 2023* |
| 75 | + |
| 76 | +Our earlier work studied how self-supervised representations can support label-efficient knowledge transfer across domains, forming part of my broader interest in robust learning under limited supervision. |
| 77 | + |
| 78 | +[\[Paper\]](https://ieeexplore.ieee.org/document/10222142) |
| 79 | + |
| 80 | +## Beyond research |
| 81 | + |
| 82 | +I enjoy being involved in the research community beyond my own projects. I have served as a reviewer for **IJCAI-ECAI** and **IJCNN**, and I have been involved in mentoring students through the **MINERVA Mentoring Program at the University of Bonn**. |
| 83 | + |
| 84 | +I am always happy to talk about reliable language models, small models, abstention, unusual model behaviours, or research ideas somewhere between *"this probably should not work"* and *"why does this actually work?"* |
| 85 | + |
| 86 | +## Contact |
| 87 | + |
| 88 | +If you would like to get in touch, feel free to email me at |
| 89 | + |
| 90 | +**mchopra[at]uni-bonn.de**. |
0 commit comments