Covary: Whole-Genome Intelligence for Agricultural Genomics
An alignment-free, translation-aware machine-learning framework for large-scale biological sequence analysis — powering phylogenomics, pathogen surveillance, and sequence-based analysis at scale for agriculture.
Fig. Covary outputs: PCA · t-SNE · UMAP · dendrograms · heatmaps — all alignment-free
What Is Covary?
Covary is a whole-genome intelligence framework — not a single-tool phylogenetic package. It is a machine-learning substrate for large-scale biological sequence analysis, powered by TIPs-VF (Translator-Interpreter Pre-seeding for Variable-length Fragments). It turns any set of sequences into a computable, clusterable, visualizable representation — without a single multiple-sequence alignment.
No MSA required
Circumvents computationally expensive multiple sequence alignments, enabling scalable analyses for large datasets — from dozens to thousands of sequences in a single run.
Codon-bound context
Incorporates codon-bound information for biologically meaningful sequence comparison — works in both coding and non-coding sequences, preserving the signal that alignment-based tools recover through expensive preprocessing.
Embedding → distance → cluster
Computes embeddings and distance matrices for downstream clustering and visualization across PCA, t-SNE, and UMAP — the embedding is the intelligence substrate for all downstream tasks.
Species to order level
Resolves sequences at species, genus, family, and order levels. Supports multi-FASTA files from a variety of organisms — the same framework across taxa.
The Covary Approach: From FASTA to Whole-Genome Intelligence
Six steps from raw sequence files to a computable intelligence substrate — no alignment, no substitution model, no reference phylogeny required.
The Pipeline in Action
Cluster Structure, Learned End-to-End
High-Resolution Similarity at Genome Scale
Scientific Validation: Benchmarked Against Gold-Standard Pipelines
Covary v1.3 has been benchmarked against established phylogenetic pipelines — ETE3, IQ-TREE, MAFFT, and FastTree — across four use cases: classification, identification, relationship, and prediction. The framework occupies a distinct computational niche: fast, alignment-free exploratory analysis at scale, complementing — not replacing — gold-standard tools for publication-grade phylogenies.
Covary performs taxonomic classification similarly with alignment-based algorithms.

Covary has similar efficiency as conventional pipelines in sequence identification.

Covary outperforms conventional pipelines in providing sequence relationships.

Covary features bring superior predictive context not found in conventional phylogenetics.

Agricultural Use Cases: Phylogenomics, Pathogen Surveillance, and Sequence-Based Analysis
Covary was built for sequence-based analysis that requires throughput, speed, efficiency, and AI inference — and agricultural genomics is a direct application domain. These four use cases map directly onto the problems Philippine and SE Asian agriculture face.
Use Case 1 — Phylogenomics for Agriculture
Use Case 2 — Pathogen Surveillance
Use Case 3 — Sequence-Based Analysis at Scale: Throughput, Speed, Efficiency, and AI Inference
Use Case 4 — Agricultural Genomics (Directly in the Covary Research Topics Library)
The Covary Platform: Three Open-Source Toolkits
Three composable, open-source toolkits — available on GitHub — make up the Covary platform. Together they form a reusable pipeline from FASTA to embedding to insight.
Covary-encoder
A k-mer-derived, non-overlapping, and frequency-independent encoding logic. Represents genetic sequences based on relative proximity, directional alignment, and translation awareness. The core encoding engine behind Covary's alignment-free approach.
Explore on GitHub →Seed Aligner
A computationally-optimized tool that detects a common seed region across genetic sequences and reorders them to start at the same point, standardizing FASTA inputs for Covary without full MSA. A pre-processing step that preserves alignment-free speed.
Explore on GitHub →Mutagen-PX
A lightweight Python toolkit that simulates tumor-specific gene sequence profiles by applying patient mutation data from TCGA cohorts to a reference sequence. Recreates mutated FASTA outputs per patient — demonstrating the framework's flexibility for generating sequence variants at scale.
Explore on GitHub →
Beyond Phylogenetics: A General-Purpose Sequence Intelligence Engine
Covary's translation-aware, alignment-free framework is not limited to traditional phylogenetics. The same substrate powers oncogenomics, epidemiology, forensic genetics, and precision medicine — the framework generalizes; only the biological question changes.
Tumor clonal evolution
Map subclonal architecture and mutational trajectories in cancer genomes — primary-to-metastasis divergence, treatment-resistance tracing. Apply Mutagen-PX to generate patient-specific mutated FASTA profiles, then embed and cluster to reveal subclonal groupings.
Outbreak & variant prediction
Model evolutionary trajectories of viral genomes in outbreak settings — forecasting emergent variants and resistance evolution for proactive public health response. Analyzed 906 SARS-CoV-2 genomes in a single run; identifies transmission clusters without MSA.
Forensic & environmental genetics
Species-of-origin determination from complex biological samples; wildlife DNA barcoding from eDNA; forensic FASTA database integration. Apply Covary's identification workflow to match unknown sequences to reference clades without alignment.
Treatment response
Sequence-level stratification of patients or pathogens; identify sequence-based patient subgroups that predict treatment response; enable rapid reclassification as new sequences or resistance mutations emerge.
Published Research
Two peer-reviewed publications document Covary's method and its application to outbreak-scale viral genomics. The underlying encoding framework — TIPs-VF — is also published.
01 — Core method
Covary: A translation-aware framework for alignment-free phylogenetics using machine learning
bioRxiv · 2025-11 · De los Santos, 2025
doi.org/10.1101/2025.11.13.687960 →02 — Outbreak application
Rapid Phylogenomic Analysis of Thousands Outbreak-Causing Viral Genomes Using Covary
Preprints · 2025-12
doi.org/10.20944/preprints202512.1970.v1 →Underlying encoding framework:
What Covary Is — and Is Not
Honest framing matters for agricultural researchers deciding whether Covary belongs in their workflow.
What We Are Seeking at Agrinnovation 2026
Covary is a deployed whole-genome intelligence framework ready for agricultural partnerships. We are looking for UPLB collaborators to extend the platform into Philippine crop and livestock pathogen genomics.
UPLB institute collaborations
Partnerships with IB, IPB, CVM, CEA, and other UPLB institutes on crop and livestock pathogen genomics — encode, cluster, and analyze Philippine-relevant pathogens together.
Philippine crop pathogen profiling
Joint phylogenomic profiling of Philippine crop pathogens — Phytophthora, Fusarium, Xanthomonas, Ralstonia, viral pathogens of rice, coconut, banana, and abaca — alignment-free, at scale.
Pathogen surveillance pipelines
Pathogen surveillance pipelines for Philippine agricultural biosecurity — rapid, alignment-free identification and transmission-chain reconstruction from field isolate sequences.
Soil & rhizosphere microbiomes
Metagenomic exploration of Philippine soil, rhizosphere, and plant-associated microbiomes — taxonomic composition, rare-organism detection, functional divergence.
Antimicrobial resistance tracking
Antimicrobial resistance tracking in Philippine agricultural bacterial populations — identify emerging resistance clades, predict spread patterns in clinical or environmental reservoirs.
Local genome access
Access to locally sequenced pathogen and germplasm genomes — draft or reference — to extend the Covary training and validation corpus with Philippine-relevant diversity.
Let's talk whole-genome intelligence for your crop pathogen, livestock disease agent, or metagenomic sample.
Covary is a deployed framework — encode once, analyze many ways. We are at AgrInnovation 2026, UPLB, 19–20 October 2026.
marvin@chordexbio.com inquiry@chordexbio.com