RNAprotCI


What is RNAProtCI?

RNAprotCI (RNA-protein Core Interaction) is a database of protein-RNA interactions integrating structural and RNA binding motif data. It provides a dynamic visualisation of the interface structures and binding motifs, as well as pre-computed agreement scores between both information sources.

It was developed by Fauconnet et al.

This work was supported by Agence Nationale de la Recherche (ANR, grants ANR-18-CE45-0005 ESPRINet, ANR-22-CE44-0044 DeepTransfo and ANR-24-CE35-2173 CDefenseRNA).

Construction of RNAprotCI

Firstly, structural data were extracted from the Protein Data Bank (PDB) [1] and structures were filtered to remove peptides (<50 amino acids) and small interaction interfaces (less than 2 interface nucleotides and/or less than 5 interface amino acids). Pre-processed RNA binding motif data were extracted from Kuret et al [2]. These data were filtered to keep cases with at least RNA compete (RNAC) or RNA Bind-n-Seq (RBNS) motif data, then matched with structural data based on UniProt protein annotations and BLAST sequence similarity search. Finally, agreement scoring between structural and omics protein-RNA interaction data was assessed via a PWM score (see "Agreement scoring" below).

Protein-RNA interface annotation

We annotated protein-RNA interfaces based on:

Agreement scoring

agreement score pipeline

The agreement score 𝑆, inspired from a transcription factor binding site predictor using position weight matrices (Siddharthan [5]), is computed by aligning the frequency matrix (from RNA binding frequencies) to subparts of the interacting RNA sequence (from the protein-RNA interaction structure) using an iterative sliding window strategy. We use dinucleotide-based scoring to better capture binding preferences. Each element 𝐷 is defined as the likelihood to observe the dinucleotide 𝑥 at position 𝑝 in the RNA sequence subpart with respect to the frequency matrix and 5-mer alignment. At a given dinucleotide position 𝑝, 𝑛𝑥𝑝 represents the number of observed dinucleotides 𝑥 in the RNA binding frequency alignment and 𝑛 represents the total number of dinucleotides at this position. A pseudocount is added to manage dinucleotides with few occurrences: 𝑓1 and 𝑓2 represent normalized frequencies of both nucleotides in 𝑥 in the aligned frequency matrix, and there are 16 possible dinucleotides. This score is applied to all dinucleotides at all positions in the RNA sequence and averaged to simplify comparisons between proteins.

List of all proteins studied in this database

rnaprotci protein list

This page provides an overview of all 63 proteins for which we were able to associate protein-RNA interaction structures and RNA binding motif data for at least RNAC or RBNS. Users can search for specific keywords within the Uniprot ID (ex: P35637), protein name (ex: FUS) or description fields, or filter to select proteins with specific RNA binding motif experiment data (RNAC, RBNS and/or CLIP) displayed as WebLogos [6]. Users can additionnaly, filter by PDB id using the dedicated search field. The β€œexplore” button directs to the table of associated structures and related information, while clicking on the UniProt ID redirects to the corresponding Uniprot database entry [7].

List of all structures of a given protein in the PDB

rnaprotci protein

This particular focus on a given protein lists all related protein-RNA structures that have associated RNA binding motif data. The PDB accession is composed of the PDB identifier, the protein chain and RNA chain (PDBid_protchain_RNAchain). Users can click on it to reach the 3D visualisation web interface. The %id represents the percentage of sequence identity between the protein sequence in the PDB structure and the Uniprot protein sequence. The right panel allows users to download the RNA binding motif alignment used to create the WebLogo, for available data sources. Above the table, two dropdown menus allow the user to pick a pair of PDB accessions in order to structurally align them.

Single structure: Visualisation of {PDBid_protchain_RNAchain}

rnaprotci single protein

The top left panel allows for the exploration of the 3D structure of an interface. In the interactive 3D view, various modes of representation can be chosen for each chain (cartoon, licorice or surface, with chain, heteroatom or chainbow coloring). The protein chain can be colored according to hydrophobicity, electrostatics, accessibility (core and rim regions) and evolutionary conservation computed with the Rate4Site program [8].
The bottom left panel provides the agreement score and the 4 nucleotides that fit the best with the selected RNA binding motif data source. Clicking on these nucleotides highlights them in the interactive 3D view.
On the right side, users can choose the RNA binding motif data source they want to focus on and the motif logo is shown. Below the motif logo, the best-fitting part of the structural interacting RNA is shown with associated dinucleotide scores as a histogram. Higher scores (close to 0) represent the best agreement scores.

Aligned structures

rnaprotci pair protein

In the case of a pair of interfaces, the aligned interface structures can also be explored interactively. The interactive 3D view behaves similarly to the 3D view of single interfaces.
In this section, structures are aligned together using US-align [9]. Agreement scores are represented for both protein-RNA interactions with two histograms.

References

This work: please cite Fauconnet et al. 2025

Tools used for dataset construction

Tools used for the webserver

This site was generated using Django and django modules. Tables are displayed using Ajax Datatables and histogrammes are plotted thanks to Chartjs. Protein-RNA interface structures are displayed in three-dimensions using the WebGL-based NGL Viewer plugin. Many thanks also to the RPBS facility and their NGL Viewer code from which we inspired ourselves to create the protein-rna interface visualisations in NGL Viewer.

Contact information

For help or to report a problem, please email us at jessica.andreani@i2bc.paris-saclay.fr (principal coordinator) and/or contact-bioi2@i2bc.paris-saclay.fr (BIOI2 bioinformatics facility).