ChordexBio UPLB Agrinnovation 2026 Back to home
Research Deck · ChordexBio at UPLB Agrinnovation Marketplace 2026
Whole-Genome Intelligence · Agriculture

Whole-Genome Intelligence for Agricultural Genomics

Marvin De los Santos — ChordexBio Covary UPLB Agrinnovation Marketplace 2026

Covary: Whole-Genome Intelligence for Agricultural Genomics

An alignment-free, translation-aware machine-learning framework for large-scale biological sequence analysis — powering phylogenomics, pathogen surveillance, and sequence-based analysis at scale for agriculture.

Agrinnovation 2026 · UPLB 19–20 October 2026 ChordexBio · Ilocos Norte
906
SARS-CoV-2 genomes in one run
4 use cases
classification · identification · relationship · prediction
vs ETE3 · IQ-TREE · MAFFT · FastTree
benchmarked against gold-standard pipelines
2 publications
bioRxiv + Preprints · 2025
Covary feature overview — PCA, t-SNE, UMAP embeddings and dendrogram outputs Fig. Covary outputs: PCA · t-SNE · UMAP · dendrograms · heatmaps — all alignment-free
01

What Is Covary?

Covary is a whole-genome intelligence framework — not a single-tool phylogenetic package. It is a machine-learning substrate for large-scale biological sequence analysis, powered by TIPs-VF (Translator-Interpreter Pre-seeding for Variable-length Fragments). It turns any set of sequences into a computable, clusterable, visualizable representation — without a single multiple-sequence alignment.

Alignment-free

No MSA required

Circumvents computationally expensive multiple sequence alignments, enabling scalable analyses for large datasets — from dozens to thousands of sequences in a single run.

Translation-aware

Codon-bound context

Incorporates codon-bound information for biologically meaningful sequence comparison — works in both coding and non-coding sequences, preserving the signal that alignment-based tools recover through expensive preprocessing.

Distance-based

Embedding → distance → cluster

Computes embeddings and distance matrices for downstream clustering and visualization across PCA, t-SNE, and UMAP — the embedding is the intelligence substrate for all downstream tasks.

Phylogenetic resolution

Species to order level

Resolves sequences at species, genus, family, and order levels. Supports multi-FASTA files from a variety of organisms — the same framework across taxa.

02

The Covary Approach: From FASTA to Whole-Genome Intelligence

Six steps from raw sequence files to a computable intelligence substrate — no alignment, no substitution model, no reference phylogeny required.

  • 1. Input: multi-FASTA files — any organism, any sequence type, coding or non-coding
  • 2. Encoding: k-mer-derived, non-overlapping, frequency-independent encoding via Covary-encoder — represents sequences by relative proximity, directional alignment, and translation awareness
  • 3. Pre-processing (optional): Seed Aligner detects a common seed region across sequences and reorders them to start at the same point — standardizing inputs without full MSA
  • 4. Embedding: sequences embedded into low-dimensional space reveal clean lineage clusters — end-to-end learned, not pre-trained
  • 5. Distance matrix: pairwise distances computed in embedding space — the basis for all downstream analysis
  • 6. Output: PCA, t-SNE, UMAP plots, heatmaps, and dendrograms — visualization and clustering without a single alignment
  • The Pipeline in Action

    Real-time progress at scale — 906 SARS-CoV-2 genomes, 16 threads

    Covary phylogenetic analysis pipeline — FASTA encoding to tree output
    Fig. Covary phylogenomic pipeline. From multi-FASTA input through k-mer encoding, distance-matrix computation, and embedding-based clustering to dendrogram, t-SNE, and UMAP outputs — all without multiple-sequence alignment.

    Cluster Structure, Learned End-to-End

    Sequences embedded into low-dimensional space reveal clean lineage clusters

    Covary t-SNE sequence embedding scatter plot — clean lineage clusters
    Fig. t-SNE embedding. Sequences projected into 2D from the learned embedding space. Distinct, non-overlapping clusters correspond to known lineages — recovered end-to-end by the model, not by pre-specified features.

    High-Resolution Similarity at Genome Scale

    UMAP distance heatmap — block structure reveals coherent sample groups

    Covary UMAP distance heatmap — block structure reveals coherent sample groups
    Fig. UMAP distance heatmap. Pairwise distances across all samples in embedding space. Dark blue blocks (on and off diagonal) = coherent clusters; red off-diagonal blocks = distinct groups. Block structure reveals sample groupings without alignment.
    03

    Scientific Validation: Benchmarked Against Gold-Standard Pipelines

    Covary v1.3 has been benchmarked against established phylogenetic pipelines — ETE3, IQ-TREE, MAFFT, and FastTree — across four use cases: classification, identification, relationship, and prediction. The framework occupies a distinct computational niche: fast, alignment-free exploratory analysis at scale, complementing — not replacing — gold-standard tools for publication-grade phylogenies.

    Covary performs taxonomic classification similarly with alignment-based algorithms.

    Workflow: Covary v1.3 (powered by TIPs-VF; De los Santos, 2025) vs. ETE3 3.1.3 (Huerta-Cepas et al., 2016)

    Method:

    1. The 16S rRNA of Staphylococcus taiwanensis strain NTUH-S172 (NR_181843.1) retrieved from NCBI Nucleotide.
    2. Up to 5 base mutations introduced, iterated 100×, creating 100 additional mutational variants. SARS-CoV-2 (NC_045512.2) as outgroup control; n = 101.
    3. Sequences fed into Covary v1.3 via Google Colab and ETE3 3.1.3 via GenomeNet.
    4. Default parameters. Covary tree via custom code; ETE3 used FastTree v2.1.8 (Price et al., 2010).
    Result: Covary v1.3 generated similar phylogenetic clustering as ETE3, delineating the outgroup from the subsequences generated from Staphylococcus taiwanensis 16S rRNA gene sequence.
    Covary v1.3 vs ETE3 — Staphylococcus taiwanensis 16S rRNA classification

    Covary has similar efficiency as conventional pipelines in sequence identification.

    Workflow: Covary v1.3 (powered by TIPs-VF; De los Santos, 2025) vs. ETE3 3.1.3 (Huerta-Cepas et al., 2016)

    Method:

    1. 16S rRNA sequences of Staphylococcus species from NCBI BioProject: txid1279[ORGN] AND (33175[Bioproject] OR 33317[Bioproject]).
    2. SARS-CoV-2 (PV950678.1) as outgroup. Unlabelled Staphylococcus spp. (NR_036791.1) and unlabelled SARS-CoV-2 (PV950681.1) as unknowns; n = 94.
    3. Sequences fed into Covary v1.3 via Google Colab and ETE3 3.1.3 via GenomeNet.
    4. ETE3 pipeline: MAFFT v6.861b alignment (Katoh et al., 2005), then ML tree via IQ-TREE 1.5.5 + ModelFinder (Nguyen et al., 2015).
    Result: Both Covary v1.3 and ETE3 identified and placed the unknown sequences in the correct clades or groups.
    Covary v1.3 vs ETE3/IQ-TREE — identification of unknown Staphylococcus and SARS-CoV-2

    Covary outperforms conventional pipelines in providing sequence relationships.

    Workflow: Covary v1.3 (powered by TIPs-VF; De los Santos, 2025) vs. ETE3 3.1.3 (Huerta-Cepas et al., 2016)

    Method:

    1. TP53 mutational profiles of patients with head and neck cancers (GDC-TCGA HNC) retrieved via UCSC Xena Browser.
    2. TP53 sequences of patients with mutations reconstructed by mapping mutation positions and types from wildtype (WT) TP53. Patients with no variant records considered WT; n = 220.
    3. Sequences fed into Covary v1.3 via Google Colab and ETE3 3.1.3 via GenomeNet.
    4. ETE3 pipeline: MAFFT v6.861b alignment, then ML tree via IQ-TREE 1.5.5 + ModelFinder.
    Result: Covary v1.3 provided better context in mutational relationship, clustering wildtypes distinctly from the mutational variants — which ETE3 failed to demonstrate.
    Covary v1.3 vs ETE3 — TP53 mutational clustering in TCGA head and neck cancer cohort

    Covary features bring superior predictive context not found in conventional phylogenetics.

    Workflow: Covary v1.3 (powered by TIPs-VF; De los Santos, 2025) — TP53 mutational profiles, GDC-TCGA HNC, n = 220

    Method:

    1. TP53 mutational profiles of patients with head and neck cancers retrieved via UCSC Xena Browser.
    2. TP53 sequences of patients with mutations reconstructed by mapping mutation positions and types from wildtype TP53. Patients with no variant records considered WT; n = 220.
    3. Sequences fed into Covary v1.3 via Google Colab implementation. Default settings.
    Result: Covary v1.3 inferred mutational distances and provided predictive genome integrity outcomes based on TP53 mutational markers.
    Covary v1.3 prediction — TP53 mutational distance heatmap with genome integrity inference
    04

    Agricultural Use Cases: Phylogenomics, Pathogen Surveillance, and Sequence-Based Analysis

    Covary was built for sequence-based analysis that requires throughput, speed, efficiency, and AI inference — and agricultural genomics is a direct application domain. These four use cases map directly onto the problems Philippine and SE Asian agriculture face.

    Use Case 1 — Phylogenomics for Agriculture

    Crop pathogens · livestock disease agents · plant pathogen diversity

  • Reconstruct evolutionary histories and phylogenetic trees from multi-FASTA sequences — resolve species relationships at multiple taxonomic levels without MSA
  • Crop pathogen strains: compare Phytophthora, Fusarium, Xanthomonas, Ralstonia, and other agricultural pathogens at scale — identify lineages, track introductions, resolve outbreak sources
  • Plant pathogen diversity: profile pathogen population structure across Philippine farm sites, seasons, and cropping systems — without waiting for reference-grade assemblies
  • Livestock disease agents: compare bacterial and viral strains affecting Philippine livestock — detection, lineage tracking, and cross-site transmission inference
  • Alignment-free advantage: when you have hundreds to thousands of sequences from field samples, MSA becomes the bottleneck — Covary's k-mer encoding is lightweight and fast
  • Covary phylogenomic pipeline for agricultural pathogen analysis
    Fig. Phylogenomic pipeline. The same Covary workflow — FASTA → encoding → embedding → distance → dendrogram — applies to agricultural pathogens at any taxonomic level, without MSA.

    Use Case 2 — Pathogen Surveillance

    Outbreak tracking · variant prediction · real-time biosecurity

  • Rapidly identify and track viral and bacterial pathogens — reconstruct outbreak transmission chains and monitor emerging variants in real time
  • 906 SARS-CoV-2 whole genomes analyzed in a single Covary run in preliminary benchmarks — identifies transmission clusters and outbreak lineages without MSA
  • Agricultural parallel: the same pipeline applies to plant viruses (Canna yellow streak virus, Rice yellow mottle virus), bacterial wilts, and fungal outbreaks — any pathogen generating genome-scale sequence data
  • Rapidly characterizes newly emerging variants and places them in evolutionary context — critical for biosecurity and precision agriculture
  • Predictive modeling of variant trajectories using embedding distance patterns — forecast which lineages may gain selective advantage
  • Covary pathogen surveillance and outbreak variant prediction
    Fig. Pathogen surveillance. Covary's alignment-free, translation-aware framework supports rapid outbreak analysis — 906 SARS-CoV-2 genomes in one run, with the same approach directly applicable to agricultural pathogens.

    Use Case 3 — Sequence-Based Analysis at Scale: Throughput, Speed, Efficiency, and AI Inference

    The four capabilities agricultural genomics actually needs

  • Throughput: Covary processes large multi-FASTA datasets without MSA — 906 SARS-CoV-2 whole genomes in a single run; scales to thousands
  • Speed: k-mer encoding is lightweight; no MSA, no substitution model optimization — minutes, not days, for exploratory analysis on large datasets
  • Efficiency: no model assumption — distance derived from embedding space rather than an explicit evolutionary model (GTR, HKY, WAG); lower computational overhead
  • AI inference: sequences embedded into a learned low-dimensional space reveal lineage clusters end-to-end — PCA, t-SNE, and UMAP projections; distance heatmaps; dendrograms — all generated from the learned representation
  • The embedding is the intelligence substrate: once sequences are encoded, the same representation supports classification, identification, relationship mapping, and prediction — not just one task
  • Covary UMAP distance heatmap — high-resolution similarity at genome scale
    Fig. High-resolution similarity at genome scale. The UMAP distance heatmap shows pairwise distances across all samples in embedding space — block structure reveals coherent groups without requiring alignment or a reference phylogeny.

    Use Case 4 — Agricultural Genomics (Directly in the Covary Research Topics Library)

    Classification · Identification · Philippine crop and livestock biosecurity

  • Agricultural genomics is one of the 9 research topics in the Covary Research Topics library — compare crop pathogen strains, plant pathogen diversity, or livestock disease agents for biosecurity and precision agriculture applications
  • Metagenomic exploration: profile mixed microbial communities from environmental or clinical samples — uncover taxonomic composition, detect rare organisms in complex datasets, analyze functional divergence
  • Antimicrobial resistance: track AMR gene evolution across bacterial populations — identify emerging resistance clades and predict spread patterns in clinical or environmental reservoirs
  • Sequence database mining: rapidly screen large public genomic databases (NCBI, GISAID) for closest relatives to query sequences — enabling database-scale comparative studies
  • All four capabilities — Classification, Identification, Relationship, Prediction — are directly applicable to agricultural sequence analysis at the scale and speed real farm and biosecurity workflows require
  • 05

    The Covary Platform: Three Open-Source Toolkits

    Three composable, open-source toolkits — available on GitHub — make up the Covary platform. Together they form a reusable pipeline from FASTA to embedding to insight.

    Encoding

    Covary-encoder

    A k-mer-derived, non-overlapping, and frequency-independent encoding logic. Represents genetic sequences based on relative proximity, directional alignment, and translation awareness. The core encoding engine behind Covary's alignment-free approach.

    Explore on GitHub →
    Pre-processing

    Seed Aligner

    A computationally-optimized tool that detects a common seed region across genetic sequences and reorders them to start at the same point, standardizing FASTA inputs for Covary without full MSA. A pre-processing step that preserves alignment-free speed.

    Explore on GitHub →
    Simulation

    Mutagen-PX

    A lightweight Python toolkit that simulates tumor-specific gene sequence profiles by applying patient mutation data from TCGA cohorts to a reference sequence. Recreates mutated FASTA outputs per patient — demonstrating the framework's flexibility for generating sequence variants at scale.

    Explore on GitHub →
    Covary-encoder toolkit screenshot
    Fig. Covary-encoder. The core encoding toolkit — k-mer-derived, non-overlapping, frequency-independent, translation-aware. The engine that makes Covary's alignment-free approach work.
    06

    Beyond Phylogenetics: A General-Purpose Sequence Intelligence Engine

    Covary's translation-aware, alignment-free framework is not limited to traditional phylogenetics. The same substrate powers oncogenomics, epidemiology, forensic genetics, and precision medicine — the framework generalizes; only the biological question changes.

    Oncogenomics

    Tumor clonal evolution

    Map subclonal architecture and mutational trajectories in cancer genomes — primary-to-metastasis divergence, treatment-resistance tracing. Apply Mutagen-PX to generate patient-specific mutated FASTA profiles, then embed and cluster to reveal subclonal groupings.

    Epidemiology

    Outbreak & variant prediction

    Model evolutionary trajectories of viral genomes in outbreak settings — forecasting emergent variants and resistance evolution for proactive public health response. Analyzed 906 SARS-CoV-2 genomes in a single run; identifies transmission clusters without MSA.

    Forensics

    Forensic & environmental genetics

    Species-of-origin determination from complex biological samples; wildlife DNA barcoding from eDNA; forensic FASTA database integration. Apply Covary's identification workflow to match unknown sequences to reference clades without alignment.

    Precision Medicine

    Treatment response

    Sequence-level stratification of patients or pathogens; identify sequence-based patient subgroups that predict treatment response; enable rapid reclassification as new sequences or resistance mutations emerge.

    Covary beyond phylogenetics — oncogenomics, epidemiology, forensics, precision medicine
    Fig. Beyond phylogenetics. The same alignment-free, translation-aware substrate powers four additional domains — the framework generalizes; only the biological question changes.
    07

    Published Research

    Two peer-reviewed publications document Covary's method and its application to outbreak-scale viral genomics. The underlying encoding framework — TIPs-VF — is also published.

    01 — Core method

    Covary: A translation-aware framework for alignment-free phylogenetics using machine learning

    bioRxiv · 2025-11 · De los Santos, 2025

    doi.org/10.1101/2025.11.13.687960 →

    02 — Outbreak application

    Rapid Phylogenomic Analysis of Thousands Outbreak-Causing Viral Genomes Using Covary

    Preprints · 2025-12

    doi.org/10.20944/preprints202512.1970.v1 →

    Underlying encoding framework:

  • TIPs-VF: Translator-Interpreter Pre-seeding for Variable-length Fragments — the encoding foundation underlying Covary. doi.org/10.1101/2025.02.15.637782
  • Covary in STR profiling: Forensic genetics protocol — alignment-free species identification from STR data. dx.doi.org/10.17504/protocols.io.ewov1ky2pgr2/v1
  • 08

    What Covary Is — and Is Not

    Honest framing matters for agricultural researchers deciding whether Covary belongs in their workflow.

  • Covary is: a fast, alignment-free, translation-aware framework for exploratory analysis of large sequence datasets — classification, identification, relationship mapping, and prediction
  • Covary is: an AI-powered whole-genome intelligence substrate — learned embeddings that generalize across classification, identification, relationship, and prediction tasks
  • Covary is NOT: a direct replacement for MEGA, IQ-TREE, RAxML, or FastTree — for high-confidence species trees with statistical support and publication-grade phylogeny reconstruction, alignment-based tools remain the gold standard
  • Covary's strength: speed and scale at the exploratory stage — rapidly screening large, heterogeneous datasets and generating embedding-based cluster hypotheses
  • Best practice: use Covary for rapid exploratory analysis and hypothesis generation across large datasets, then follow up with targeted alignment-based analyses on subsets of interest
  • For the agricultural researcher: Covary accelerates the front end of your phylogenomic and pathogen-surveillance workflow — it is a complement to, not a substitute for, gold-standard tools
  • Covary outperforms ETE3 on sequence relationship context — TP53 mutational clustering
    Fig. Honest validation. Covary outperformed ETE3 on relationship context (TP53, n=220) — but the honest claim is that it complements gold-standard tools, not replaces them. The same posture applies to agricultural use.
    09

    What We Are Seeking at Agrinnovation 2026

    Covary is a deployed whole-genome intelligence framework ready for agricultural partnerships. We are looking for UPLB collaborators to extend the platform into Philippine crop and livestock pathogen genomics.

    Partners

    UPLB institute collaborations

    Partnerships with IB, IPB, CVM, CEA, and other UPLB institutes on crop and livestock pathogen genomics — encode, cluster, and analyze Philippine-relevant pathogens together.

    Phylogenomics

    Philippine crop pathogen profiling

    Joint phylogenomic profiling of Philippine crop pathogens — Phytophthora, Fusarium, Xanthomonas, Ralstonia, viral pathogens of rice, coconut, banana, and abaca — alignment-free, at scale.

    Surveillance

    Pathogen surveillance pipelines

    Pathogen surveillance pipelines for Philippine agricultural biosecurity — rapid, alignment-free identification and transmission-chain reconstruction from field isolate sequences.

    Metagenomics

    Soil & rhizosphere microbiomes

    Metagenomic exploration of Philippine soil, rhizosphere, and plant-associated microbiomes — taxonomic composition, rare-organism detection, functional divergence.

    AMR

    Antimicrobial resistance tracking

    Antimicrobial resistance tracking in Philippine agricultural bacterial populations — identify emerging resistance clades, predict spread patterns in clinical or environmental reservoirs.

    Data

    Local genome access

    Access to locally sequenced pathogen and germplasm genomes — draft or reference — to extend the Covary training and validation corpus with Philippine-relevant diversity.

    Let's talk whole-genome intelligence for your crop pathogen, livestock disease agent, or metagenomic sample.

    Covary is a deployed framework — encode once, analyze many ways. We are at AgrInnovation 2026, UPLB, 19–20 October 2026.

    marvin@chordexbio.com inquiry@chordexbio.com