Descriptor-Based Pre-Training Improves Reaction Property Prediction
Published in ChemRxiv (preprint), 2026
Recommended citation: Akshat Shirish Zalte, Cheng Fang, Yiting Zheng, Essam Metwally, Alan Cheng, Alec Glisman, and Hao-Wei Pang. (2026). "Descriptor-Based Pre-Training Improves Reaction Property Prediction." ChemRxiv. doi:10.26434/chemrxiv.15008692/v1. https://doi.org/10.26434/chemrxiv.15008692/v1
Abstract
Pre-training graph neural networks on large unlabeled datasets has substantially improved molecular property prediction, yet pre-training approaches for reaction property prediction remain underexplored. We introduce CheMeleon-Rxn, a pre-trained reaction graph neural network that adapts the CheMeleon framework to condensed graphs of reaction (CGR) and demonstrates that descriptor-regression is an effective pre-training strategy for low-data reaction property prediction. An O(10M)-parameter D-MPNN encoder was pre-trained on ~2 million reactions by regressing onto dense reaction descriptors pooled from classical molecular descriptors, then fine-tuned on small reaction property prediction datasets. Across regression tasks spanning gas-phase activation energies, reaction enthalpies, experimental yields, and rate coefficients, CheMeleon-Rxn is best or statistically tied-for-best on seven of eight tasks, outperforming the same encoder trained from scratch as well as classical-descriptor, fingerprint, and pre-trained-fingerprint baselines. The pre-training strategy also generalizes to other edge-aware graph neural network backbones and is robust to the choice of descriptor set, pooling operation, and other pre-training hyperparameters.
| Download paper here | Code (fork) |
Recommended citation: Akshat Shirish Zalte, Cheng Fang, Yiting Zheng, Essam Metwally, Alan Cheng, Alec Glisman, and Hao-Wei Pang. (2026). “Descriptor-Based Pre-Training Improves Reaction Property Prediction.” ChemRxiv. doi:10.26434/chemrxiv.15008692/v1.
TOC Graphic

CheMeleon-Rxn pre-trains a condensed-graph-of-reaction D-MPNN encoder to regress Mordred-Rxn descriptors on ~2 million reactions, then fine-tunes the encoder on small labeled reaction property datasets.
