Mentee Testimonials

Full testimonials written by students I have mentored. Names, paper titles, venues, programs, and other identifying details have been removed or generalized to preserve the authors' anonymity.

I have worked with Thomas for nearly a year across three projects in mechanistic interpretability and AI safety. Our first project explored the development of more robust methods for understanding how fine-tuning changes a model's internals. This work resulted in a paper accepted at a workshop at a major ML conference and is currently under review at another venue. Our two ongoing projects investigate further questions in AI safety and interpretability, with the intention of submitting both to a top ML conference.

Thomas and I meet weekly for approximately one hour to review results, diagnose problems, and determine the next experiments. He is also exceptionally responsive between meetings. When I share results or ask questions, he typically responds within an hour. What distinguishes his mentorship is his strong technical intuition, particularly in mechanistic interpretability. He can often recognize which directions are genuinely promising before we invest weeks in experiments, while still encouraging me to test ideas and reach my own conclusions.

One example arose during one of our ongoing projects. I was initially unconvinced that an unconventional approach was worth pursuing, but Thomas recognized its potential and encouraged me to investigate it further. After reviewing the relevant literature and conducting initial experiments, I found that the approach worked remarkably well. These findings have since developed into a promising research direction, and we plan to publish an initial write-up soon. This experience illustrates how Thomas challenges my assumptions while giving me the space to evaluate ideas independently.

Thomas has also helped me address an important limitation in my research approach. I have generally been comfortable developing new methods and architectures that produce strong empirical results, but I have found it more challenging to analyze those results deeply and construct a compelling research narrative around them. Through his feedback, Thomas has taught me to ask sharper questions, distinguish observations from explanations, design experiments that directly test the central claim, and build stronger evidence-based arguments. I am still developing these skills, but I have already made meaningful progress.

More broadly, working with Thomas has made me substantially more independent as a researcher. I have become better at generating my own research ideas, assessing their novelty and feasibility, and taking greater ownership of projects, from initial formulation and experimentation to analysis and paper development. Without his mentorship, I believe I would have remained more focused on achieving strong numerical results rather than identifying and investigating the deeper scientific questions behind them.

Our collaboration has already produced several concrete outcomes: a workshop paper at a major ML conference, a submission under review at another venue, two ongoing projects targeting a top ML conference, a forthcoming public write-up, and a grant proposal. I also believe that our broader research agenda — developing methods to identify, interpret, and elicit hidden model behaviors — is highly promising. As increasingly capable models may exhibit consequential behaviors that do not appear under standard evaluations, this research addresses an important gap in AI safety and mechanistic interpretability.

Overall, Thomas has been one of the most thoughtful and impactful research mentors I have worked with. His combination of technical insight, responsiveness, and commitment to developing researchers' independence has substantially shaped my growth. I strongly recommend him as a mentor and research leader, and I believe his research agenda has significant potential to advance AI safety and mechanistic interpretability.

— Mentee

I first met Thomas at a workshop at a major ML conference. We hit it off right away and found that we agreed on a lot, including the importance of AI safety. We kept crossing paths at the workshop party and then at the after-party. By the end of the night, he had told me he would be up for collaborating if I was interested. Once I got home, I sat on the offer. He was far more senior than me, and I was clearly the one with more to gain. Frankly, I lacked the confidence to accept. Taking him up on it felt like forcing him into an obligation to supervise me. Three days later, before I had decided anything, Thomas messaged me to ask whether I had made it home okay. A project proposal was already attached. We have worked together ever since.

The first proposal did not stick, and neither did the next few. Thomas never treated that as wasted time. Whenever we hit a dead end, we were brainstorming again within days. From the beginning, he asked what I wanted from the collaboration and what kinds of problems I wanted to work on. Once I chose an area, he sent a shortlist that separated the papers I needed to read from the directions worth exploring. He even flagged work he had not gotten to himself but suspected would be useful. None of it came with pressure to settle on a project by the end of the meeting. He would rather spend time at the start understanding why a question matters and whether it is the right one to pursue than commit too quickly and scramble later. I had already learned the hard way, on my first workshop paper, what it costs to skip that step. What surprised me was how much energy Thomas put into teaching me that process rather than just getting us to a result.

He also picked up quickly on how I work. I like to explore a problem top down and form my own view before going bottom up to test hypotheses. Once he saw that, he stopped assigning me experiments and started bringing me directions instead. He would walk me through the evidence and ask what I thought. Then we would debate it. His instincts were usually better than mine, but he asked anyway. When I was wrong, he made sure I understood why instead of overruling me. When I got stuck, he helped me find the thread again without taking over. The project remained mine to think through, not merely mine to execute.

The project was in AI safety, asking how a recently observed failure mode arises and whether it can be reduced during training. A few months later, we had two papers accepted at a workshop at a top ML conference, with Thomas as senior author and me as co-first author.

English is not my first language, and as we wrote those papers, Thomas's advice went well beyond correcting my prose. He taught me to think of a paper as a door into the problem for someone encountering it for the first time. One figure in one of them looked clear to me because I already knew what it meant. Thomas began by telling me what worked. Then he pointed out that the figure mixed quantitative and qualitative information in a way someone deep in the field would follow, but a reader arriving through the paper would not. I rebuilt it, and it became much clearer.

He is generous with his time, too. Our weekly meetings are booked for an hour and often run past ninety minutes because we are still brainstorming. He is quick to reply, and that was true even early on, when the results I sent him were a mess. When another collaborator joined us, I saw that this was not a one-off. Thomas gave them the same attention but not the same guidance. He adapted to how each of us thought and worked. His mentorship also extends beyond the project. I can book time with him for career questions or anything else I need to think through. If I mention a paper I found interesting, he often knows one of the authors and offers an introduction. The simplest way I can put it is that he actually cares about my success. Knowing that gives me more energy to put into the work.

It was Thomas who pointed me toward research programs I would not have found on my own. He helped me polish my CV and decide which ones genuinely fit my interests. He knows that ecosystem far better than I do, and he was candid about which mentors and programs would suit me and which would not. Today, I am a research fellow at one of those programs, and I have reached the final stage of another.

Beyond the papers and the programs, the larger change is in how I do research. When we started working together, I knew how to run experiments, but I did not yet know how to scope a project with this much independence or how to make a decision like withdrawing a submission when the result no longer held. Now I do. Thomas never made those decisions for me. He taught me how to make them. Accomplishments and technical ability aside, he is a genuinely good person who cares about the people he mentors. Time and again, he has helped me in ways I would not have known to ask for. I hope to keep working with him for a long time, and I already envy his future mentees.

— Mentee

I have been working with Thomas throughout my Ph.D. program. Our primary research area is mechanistic interpretability, with a particular focus on understanding and mitigating undesirable model behaviors. This work has resulted in two papers accepted at a workshop at a top ML conference.

Our current project explores a new question about what shapes the behavior of large language models, which we aim to submit to a top ML conference.

This is my first experience having a research mentor, and it has been extremely positive. Thomas and I meet twice a week (once for an hour and once for 30 minutes) and we also communicate daily to discuss the project and share results. He has always been exceptionally responsive.

What I have appreciated most about this mentorship is how much Thomas has helped me gain confidence in myself as a junior researcher. Rather than simply giving me the answers, which would probably have been easier for him, he always encouraged me to think through the problem first, organize my ideas, and explain my reasoning. He would then share his perspective and help me improve my approach. He is not only an excellent researcher but also highly skilled at helping others grow into independent researchers.

One concrete example comes from one of our papers. The initial goal of the project was to identify a specific mechanism in the model's internal representations. However, the project did not develop in the direction we initially expected, and I became somewhat uncertain about how to proceed. Thomas helped me narrow the scope of the project and identify a clearer and more meaningful contribution. This was an especially valuable learning experience for me. Research is a field in which we often encounter more failures than successes, and Thomas taught me how to navigate these situations constructively, helping me develop the mindset and resilience required to become a strong researcher.

I also had limited access to computational resources and services such as APIs for LLM-based evaluation. Thomas made sure that I had an appropriate research environment by providing access to a powerful GPU for more than five months, as well as API credits. This support made a huge difference to my work. Without these resources, it would have been very difficult to run the experiments needed for our projects, and I do not think we could have completed them in the same way. I am especially grateful that he recognized these limitations and took concrete steps to make sure they did not hold back my research.

Thomas has also supported me beyond our immediate research projects. He has taken an active interest in ensuring that my Ph.D. progresses smoothly, and he has helped me identify internship and fellowship opportunities. He provided me with a list of relevant fellowships, offered to refer me, and has consistently been willing to introduce me to other researchers and help me develop new collaborations.

Overall, I could not have wished for a better mentor. His expertise in AI safety and mechanistic interpretability, and his genuine commitment to helping others make him an exceptional mentor for junior researchers. I firmly believe that he can make a meaningful and lasting difference in the lives and careers of many other researchers.

— Mentee