In the last decade, adoption of exome and genome sequencing has become the foundation of genomic medicine, improving diagnosis of rare disease and targeted management of some cancers. Technological advances are changing the way geneticists view genome sequencing, with advances opening new possibilities and continuing to push our understanding. Below, I explain the technological transitions in genomics, and what this could mean for patients.
When whole does not mean complete
You may be familiar with the term whole genome sequencing (WGS). Instead of ordering specific genetic tests, based on the phenotypes of the patient, WGS allowed more of the genome to be interrogated, identifying more variants and increasing diagnostic yield. The use of the word ‘whole’ was reflecting that all extracted DNA in a sample was sequenced, rather than that the whole genome was being analysed. The main reason for this distinction comes from primarily using short read sequencing technologies.
Short read sequencing captures short DNA fragments, typically between 150 and 500 basepairs, depending on the method. All short read sequencing depends on mapping reads to a reference sequence, using bioinformatics to predict where a sequence most likely fits. While this method detects most known pathogenic variants, these reads can be too short to identify large structural variants or to accurately analyse more complex regions of the genome. Some regions of the genome are known to be difficult to sequence, such as highly repetitive regions or where there is a large repeat sequence, such as pseudogenes. This introduces error, preventing detection of some variants, and this can miss diagnoses.
Currently, around 25% to 50% of patients receiving WGS will receive a diagnosis (known as diagnostic yield), depending on the condition. There is growing evidence that part of this ‘diagnostic gap’ is driven by technological limitations.
Achieving (near) perfect sequencing
Scientists are exploring a technological convergence that they propose could lead to near-perfect sequencing. This centres on three core technological advances shifting what is possible in genomic medicine.
- Longer sequences
We at PHG described the opportunities of long-read sequencing (LRS) and potential clinical applications in previous briefings. Long-reads span larger regions of the genome resolving complex regions and improving detection of structural variants. Additionally, technologies directly measure the DNA strand, capturing DNA modifications, which have a known role in some genetic diseases. The evidence that LRS can boost diagnostic yield is growing. Fundamentally, LRS overcomes key limitations of short read sequencing outlined above.
This is not to say that long read sequencing does not still have limitations, particularly lower throughput of sequencing, compared to short read sequencing, and it can still struggle to sequence or detect some types of variants. The greater challenge has been developing and validating bioinformatic tools used by clinical scientists to interpret variants in a diagnostic setting. This is particularly challenging where there is no existing benchmark, because these technologies are identifying variants that previously were missed.
There is a growing interest in clinical implementation of long reads. In the UK, NHS Genomic Networks of Excellence are exploring the opportunities of long read sequencing, as well as seeking to understand current limitations, and this could pave the way for implementation.
- Reflecting the true genome
The current reference genome is known to be limited and a pangenome aims to capture the spectrum of human diversity. A pangenome is the complete set of sequences from multiple individuals and this approach allows us to better capture the breadth of genomic variation. Advances in the pangenome have been shown to reduce errors in variant detection and improve detection of structural variants. The main difference between the current reference and the pangenome is that it moves away from seeing the genome as a linear sequence and actually seeks to reflect the realities of human diversity.
While underway, challenges remain for selecting the most appropriate “path” based on an individual’s ancestry. These efforts could be further enhanced by global efforts to enhance population genomics and develop population specific telomere-to-telomere genomes. These advances have the potential to improve the equity of genomics, leading to more accurate diagnoses for currently underserved populations. Using a genome based on population ancestry has been shown to reduce errors when mapping sequence data and to enhance variant detection for patients.
Combined with long-read sequencing, we can now better ‘phase’ variants to say which chromosome they were inherited from (remembering that we each inherit two of each 23 chromosomes, including either XX or XY). Phasing allows clinical scientists to understand variants in context and this can be incredibly useful when understanding if the variant causes disease.
With methylation data, it can even be possible to determine which parent the variant was inherited from and in the future, we may not need to sequence the parent. This could make genomics more equitable, because it is not always possible to sequence the ideal “trio” of the patient and both parents. This would improve the chance of a diagnosis for all patients, where currently there is a known diagnostic gap. On the other hand, this may infer information about the patient’s parents, who have not been consented for the test. This raises ethical challenges around an individual’s ‘right not to know’ and for health services they may feel a responsibility to contact without insight into parent wishes.
- AI to address variant ‘abundance’
Advances in genomics are coming with a side-effect: a growing list of variants where we don’t know if they cause disease, known as variants of significance (VUS). Currently, genome sequencing using short read sequencing identifies between 4 and 6 million variants per individual for all variant types. Adoption of long-read sequencing and pangenomes is likely to exacerbate this problem, particularly if it improves detection of variant types that are less well understood.
It is unsurprising then that the genomics community is keen to build on existing bioinformatic tools using machine learning and AI. There is a clear need for scalable tools to manage this interpretation bottleneck, improving interoperability of different databases and supporting variant classification. AI could help in clinical pathways, for example by extracting phenotypes in electronic health data and compiling variant evidence to reduce manual curation work by scientists.
AI cannot replace multidisciplinary involvement, and insights from the patient and family, but it could enable clinical genomics professionals to come to a diagnosis more efficiently. The real challenge will be validating these tools, particularly with concerns around hallucinations and transparency of how AI came to its ‘answer’, to ensure variant classification continues to achieve the same standards in terms of consistency and rigour.
Why ‘near’, not perfect?
So far, we have set out that genomics is in the next phase, even transition, as these technologies mature. The offering is attractive: streamlining test pathways, addressing variant abundance from technologies that arguably tell us too much, and the opportunity for greater equity by better representing diversity in the tests we perform.
As ever, these things are never simple. Other technologies will be needed, for example RNA sequencing can more accurately tell us what a variant does rather than relying on prediction, known as functional evidence. In reality, gaps and complexity will continue to remain. The future for genomic medicine is exciting but what does this mean for patients now? There is growing understanding between the offering of WGS and the realities of what can be identified. However, this has not been reflected in how we communicate this test to patients.
Consent is key
Pre-test counselling for WGS with clinicians is designed to provide informed consent and is the backbone of testing. Clinicians from across the NHS, without a background in genomics, are increasingly referring patients to have WGS and are now in the position of explaining ‘whole’ genome sequencing.
There remains an active debate about what is required for ‘informed’ consent. In the NHS, this conversation focuses on the information patients will need when they receive a test result: possible test outcomes, implications of testing for other family members, how genomic data is stored and insurance implications. Consent may include the advantages and limitations of testing, many of which have been discussed above, such as difficult to sequence regions of the genomics. It also emphasises important implications for patients, including the increased likelihood of incidental findings or variants of uncertain significance, and that WGS results can take longer to come back than other genomics tests.
It is no wonder that patients can find consent for WGS an overwhelming experience as they seek to understand this information in the context of their lives and make an informed choice on testing. Clinicians equally may have limited experience and need appropriate support to facilitate consent conversations. Continuing to reflect and critique the WGS service is important, as we work on how to communicate the important information. Words matter and this may explain why the genomics community is becoming uneasy with the use of ‘whole.’ The use of the word can suggest to a patient complete and thorough analysis. It is worth asking the question of if the additional explanation of the limitations of testing is a technical point and if it is truly helpful for patients when making the decision to receive testing.
Given these rapid and ongoing advances in genomic technologies, clinical services should revisit the language used to describe clinical genomic testing. These technologies will continue to evolve and it is important that the language reflects the changing realities of the test patients are receiving now.
