Jonas (Gus) Gustafson: 'We now have a catalogue of long read sequencing-derived structural variants we can use to filter for candidate disease-causing variants in unsolved cases.'
[EDITOR’S NOTE: Jonas (Gus) Gustafson, a Ph.D. candidate in the Miller Lab, is the first author on a paper published in the journal Bioinformatics. The paper, “needLR: long read structural variant annotation with population-scale frequency estimation,” discusses filtering and prioritization of pathogenic structural variants using data from the 1,000 Genomes Project. Here, Gustafson explains the paper’s thesis, its findings, and implications for future research.]
What is the most important aspect of this paper?
The most important aspect is to inform the scientific community that we now have a catalogue of long read sequencing-derived structural variants we can use to filter for candidate disease-causing variants in unsolved cases. Before we started doing long-read sequencing on publicly available, bio-banked cell lines from the 1,000 Genomes Project, most of the structural variants we knew about were from short-read sequencing. However, short-read sequencing identifies only about one-third of the structural variants per genome.
So, even having long-read sequencing on a patient, you could not know from those structural variants which ones are just normal in the population or could be disease causing. Now, we have long-read sequencing data on 500 samples from the 1,000 Genomes Project. With this new and open-sourced catalogue of variants, now we can say, “This variant is seen in three, or 70, or all 500 individuals and it’s probably not causing this ultra-rare disease” or “This variant has never been seen, even in the long-read data, and might be biologically relevant.”
How does this advance genome research?
The samples from the 1,000 Genomes Project are all self-reported healthy individuals. So, the idea is that any variants seen in those genomes should not be pathogenic, with the caveat that, some of them are likely carriers for recessive conditions. Further, we don’t have access to information about any phenotypes that arose after they provided the samples, so if there was a late-onset phenotype, we would not know about it. But having this publicly available data set of healthy individuals allows us to build all these control tools like needLR. We have the opportunity to understand the breadth of the structural variant landscape fully, which we could not do with short-read sequencing.
'Having this publicly available data set of healthy individuals allows us to build all these control tools like needLR.'
How did this paper come about?
When I joined the Miller Lab in the spring of 2023, the 1000 Genomes Project Long Read Sequencing Consortium – led by Drs. Danny Miller and Evan Eichler – had just started performing long-read sequencing on the samples. My responsibility became managing those samples coming in, growing the cells for DNA isolation, and then managing the bioinformatic pipeline after they were sequenced. So, I got to build that dataset from the initial 40 or 60 genomes, until we finished number 500. When I started the poster I created for my rotation talk, it was the earliest version of what later became this paper.
Who participated in the 1,000 Genomes Project?
We do not know anything about the participants other than their self-identified ancestry, among 26 world-wide populations, and biological sex. The project’s leaders aimed to capture all variants in at least 1 percent of the human population. However, there are groups that are notably missing from the dataset, including but not limited to, indigenous Australians and individuals of Middle-Eastern ancestry.
How did you come up with the name “needLR?” (pronounced “NEED-lur”)
One of my colleagues, Dr. Nikhita Damaraju, came up with the name. It is a triple entendre: “Needle in a hay stack,” a nod to Space Needle here in Seattle, and the “need” for long-read sequencing to do the work. I think it’s a great name.