Vishvak, thanks for taking the time to chat with me today. You are currently doing your PhD in Computer Science with supervisor Yue Li and co-supervisor Jun Ding. How did you end up studying computer science?
I began my undergrad at McGill in Biology, because it was my favourite course in high school. In my first year at McGill, I took the intro to computer science course, COMP 202. Initially, I found it really challenging - it was a new way of thinking that I wasn’t used to. But then midway through the course, something just clicked and I just started to enjoy it and understood what was happening and I then switched into doing joint honours in computer science and biology. My research journey began in my 4th year when I worked on two separate projects, one with each of my current supervisors. And now in grad school it has been a joint effort between me and Profs. Li and Ding.
Last year you were awarded as a D2R Doctoral Scholar. When I think of computer science, I don’t really think of RNA therapeutics research. How did you learn about D2R and what made you feel like your work would be a good fit?
I first learned about D2R through emails promoting the D2R Scholar Awards. And you're right, computer science probably isn't the first field that people associate with D2R, but I think there is a very strong connection. In modern genomics and RNA-based therapeutic development, there’s so many complex data sets available that make computational methods essential to derive insightful biological information.
Also, what appealed to me with the Scholar Awards was that it’s not just financial support, it’s also the Training Program. D2R offers exposure to trainees from so many different backgrounds, along with workshops and the annual D2R Symposium. For example, at the Demystifying Data Science workshop, there was a discussion on reproducibility that made me think about how computational results need to be able to translate to a wet lab environment.
You’ve touched on things related to your project, where you are working on developing AI-driven models to analyse single cell and spatial genomics data. Can you explain your research?
Sure. The broad goal of my project is to understand how cells change over time – how they differentiate, develop, and respond to different drugs. As you mentioned, I'm working with single cell and spatial transcriptomics data. For those that may not be familiar, single cell data allows you to measure gene expression at the individual cell level. And spatial transcriptomics provides the location of these cells, kind of mapping the transcripts to a physical location. Our goal is to use both perspectives to build an AI model that can monitor how the cell is changing over time.
Let's say you have cells sequenced at two time points and you want to see how or what's driving the change over time. The problem is when you sequence these cells, you have to destroy them, so they’re not going to be the exact same cell. My project is trying to link these cells over time. Once you have this link between them, you can look at how they've changed or what the gene programs are that drive these changes. From there, you could even try to make predictions on future changes.
What does adding in the AI element to this project allow you to do that you wouldn’t otherwise be able to achieve?
A lot of it wouldn't be possible without AI. Especially the last part that I mentioned about predicting how cells at the first time point will transform at the second time point. That's not something you can do without predictive modelling enabled by AI. Also, in single cell genomics, data can be high dimensional and noisy. AI methods can outperform other strategies to create less noisy, more accurate, lower dimensional data that can then allow for integration of, for example, gene expression with protein levels or chromatin accessibility.
What would be the end goal of this type of project – to take a snapshot of a cell and then make predictions? What are you using as your ground truth data for validation?
That’s a really important question. I’m currently using a lot of publicly available atlases, with 10,000 cells or 500,000 cells across multiple time points. So, I can use the actual cells at the second time point as the ground truth to evaluate how well we are predicting cell type transitions. So, if a cell at one time point is being linked to a cell at another time point, are these cell types the same? If it's in development, they may transition but does that transition make biological sense? There's a lot of downstream things we can do from this, some of which you could get validated by some wet lab collaborators. And that might be something we pursue in the future. But we're not there yet.
How could you see this type of work leading to the development of RNA-based therapeutics?
If we are able to predict how a cell will change, we can then introduce perturbations, for example, such as mimicking changes in gene expression. Then, using our AI model, we would see the impact on the cell and possibly identify a therapeutic target that could then be validated in a wet lab. I think the end goal of a lot of computational biology or bioinformatics research is to do just that, accelerate the identification of therapeutic targets. I really feel like D2R is working to bring together computational biologists and wet lab researchers so that collaborations can be developed at the outset of projects, where co-development of projects would be the most impactful.
Before we wrap up, is there anything about you that we haven’t touched on that you want the D2R community to know about you?
Outside the lab, I love to be as active as possible. Going to the gym, going for runs with friends. I love playing and watching sports. I sometimes joke that this is my second PhD and my first PhD is in sports!
This conversation with D2R Training Program Officer, Anthony Van Kessel, has been edited for length and clarity.