RNA inverse design · Reinforcement learning

MORAD

Reinforcement Learning with Multi-Objective Rewards for RNA Inverse Design

Jixin Tang* Ji Guo* Jun Zhao†
Fudan University
*Equal contribution  ·  †Corresponding author
MORAD training loop: group sampling, six measurements, bounded reward and group weights, forward-process update
MORAD training loop. Backbone-conditioned sampling produces candidate sequences. Three evaluation channels provide the six-component reward and group-relative weights. Forward-noised candidates train the shared policy under a frozen-reference constraint; EMA feedback refreshes the sampler.
+16.0%
GDT-TS
0.3300 → 0.3827
+19.4%
Pairing MCC
0.6110 → 0.7295
−27.5%
Normalized ensemble defect
0.3906 → 0.2830
10 / 11
Metrics improved
over a tertiary-only reward
01

Abstract

We introduce MORAD, an online reinforcement learning framework for backbone-conditioned RNA diffusion models that trains a single shared policy across diverse target backbones. Structure-conditioned RNA inverse design seeks nucleotide sequences compatible with a target three-dimensional backbone, yet existing reinforcement learning approaches either optimize a separate policy for each target backbone or rely on offline preference optimization, leaving online post-training of a reusable backbone-conditioned model largely unexplored. MORAD combines backbone-conditioned group sampling, group-relative reward normalization, and a forward-process diffusion objective. Beyond existing RL rewards that primarily target tertiary structural quality, MORAD jointly optimizes six measurements spanning tertiary geometry, base pairing, ensemble behavior, and sequence composition. To aggregate these heterogeneous objectives despite differences in scale and optimization direction, MORAD maps each measurement to a bounded desirability score and combines them with a weighted geometric mean. Compared with a tertiary-structure-only reward, this multi-objective formulation improves 10 of the 11 evaluation metrics that have a preferred direction. Applied to pretrained RIDE, MORAD improves GDT-TS by 16.0% (from 0.3300 to 0.3827) and pairing MCC by 19.4% (from 0.6110 to 0.7295), while reducing normalized ensemble defect by 27.5% (from 0.3906 to 0.2830). On the test set, MORAD outperforms RiboDiffusion, gRNAde, and RDesign across tertiary structural accuracy, secondary-structure pairing, and thermodynamic quality, establishing a new state of the art in structure-conditioned RNA inverse design.
02

Method

MORAD post-trains one backbone-conditioned diffusion policy across many target backbones, so the trained model designs sequences for a new backbone directly at inference, with no per-target optimization.

Shared-policy online post-training

Backbone-conditioned group sampling, group-relative reward normalization and a forward-process diffusion objective, applied to the pretrained RIDE model.

Six-objective reward

Tertiary geometry (GDT-TS, TM-score, RMSD), base pairing (MCC), ensemble behavior (normalized ensemble defect) and sequence composition.

Bounded desirability, geometric mean

Each measurement is mapped to a bounded desirability score and the scores are combined with a weighted geometric mean, so no single objective dominates.

The MORAD reward

RIDER and MORAD reward construction
RIDER and MORAD reward construction. RIDER scores tertiary structural consistency. MORAD combines three tertiary measurements, pairing agreement, normalized ensemble defect and composition deviation; each desirability score lies in [δ, 1].
1 · Normalize
$$u_j=\mathrm{clip}\Big(\tfrac{m_j-a_j}{b_j-a_j},\,0,\,1\Big)$$
2 · Desirability
$$d_j=\max\Big(\delta,\,\tfrac{1-\cos(\pi u_j)}{2}\Big)$$
3 · Aggregate
$$R=\exp\Big(\textstyle\sum_{j\in\mathcal J}\tilde w_j\log d_j\Big)$$

Each measurement \(m_j\) runs between a low anchor \(a_j\) and a high anchor \(b_j\); \(\tilde w_j = w_j / \sum_{k\in\mathcal J} w_k\) normalizes the weights over the available measurements \(\mathcal J\).

Measurement\(a_j\)\(b_j\)\(w_j\)
GDT-TS0.200.750.25
TM-score0.200.700.30
RMSD (Å)12.02.00.30
Secondary-structure MCC0.300.850.05
Ensemble defect / nucleotide0.600.150.05
Composition deviation0.500.020.05
03

Results

All models are evaluated on the same 153 test targets. MORAD is the policy after update 390; best values are in bold.

Candidate 0 of each target, structures predicted with RhoFold+.

ModelGDT-TS ↑TM (C4′) ↑RMSD (Å) ↓TM (C1′) ↑Success ↑Recovery ↑lDDT (C4′) ↑pLDDT ↑Coarse clash ↓
RiboDiffusion0.31320.283910.92150.36130.25490.51650.61740.693773.5383
gRNAde (T=0.1)0.35910.31798.83740.36420.24840.51150.66070.648223.7261
RDesign0.32910.290210.34410.33920.25490.47200.63680.620240.4457
RIDE0.33000.28419.71990.35380.24840.50390.66220.667122.7503
MORAD0.38270.34007.91960.38550.30070.50560.67740.672611.9438

Training dynamics

Structural scores over training updates 0 to 390
Structural-score trends over all 40 evaluations during training (updates 0–390). Thick lines show a centered five-point moving average; faint traces show the raw scores, with one sampled candidate per target.
04

Data

527
Training targets
27–258 nt · median 75
153
Test targets
15–186 nt · median 65

The test set is built separately from the training set. It comprises 71 RNA3DB entries, drawn from RNA3DB components that contain no training target, and 82 entries curated from the RCSB PDB. No test sequence exactly matches a training sequence. Both sets are released as MORAD-Targets on Hugging Face.

05

Get started

The repository contains the full training and evaluation code, configuration files for every reported run, and scripts that regenerate each table and figure of the paper. Training takes about 2 hours on 4 × RTX 3090; evaluation runs on a single GPU. Step-by-step instructions are in the README.

git clone https://github.com/Gabrile166/MORAD.git && cd MORAD
bash scripts/setup_third_party.sh       # RhoFold+, US-align, EternaFold
# create the two conda environments, see docs/INSTALL.md
bash scripts/download_assets.sh         # RIDE, RhoFold+, MORAD-RIDE, MORAD-Targets
bash scripts/evaluate_test153.sh morad  # evaluate on the 153 test targets
bash scripts/train_morad.sh             # train from pretrained RIDE (4 GPUs, ~2 h)
06

Citation

@misc{tang2026morad,
  title  = {{MORAD}: Reinforcement Learning with Multi-Objective Rewards for {RNA} Inverse Design},
  author = {Tang, Jixin and Guo, Ji and Zhao, Jun},
  year   = {2026}
}