NUMonomer enables accurate and scalable nucleic acid structure prediction from primary sequence alone
NUMonomer enables accurate and scalable nucleic acid structure prediction from primary sequence alone
si, y.; zhang, s.; chen, l.
AbstractAccurate and efficient prediction of three-dimensional nucleic acid structures can accelerate functional characterization and enable downstream applications. Recent deep-learning methods have substantially improved nucleic acid structure prediction by incorporating auxiliary inputs such as multiple sequence alignments, secondary-structure annotations, and representations from pretrained language models. However, prediction accuracy remains limited, and generating these auxiliary inputs can be computationally expensive. Here we show that learning the hierarchical organization of experimentally determined structures across multiple scales, from recurring local conformations to global fold topologies, together with exploiting representations shared between RNA and single-stranded DNA, improves model generalization. Guided by these findings, we developed NUMonomer, an end-to-end deep-learning framework trained with input sequences spanning thousands of nucleotides on a joint RNA and single-stranded DNA dataset to predict nucleic acid structures directly from sequence. Despite requiring no auxiliary inputs, NUMonomer matches or outperforms leading prediction methods on benchmarks comprising CASP16 RNA targets and non-redundant sets of experimentally determined RNA and single-stranded DNA structures, with particularly pronounced improvements for longer RNAs. Its efficient and scalable architecture also reduces inference costs by approximately two orders of magnitude relative to the evaluated methods, enabling large-scale structure prediction. Together, these findings provide insight into generalization in biomolecular structure learning and establish NUMonomer as a practical framework for nucleic acid structure prediction.