CartograPlant is a cyberinfrastructure platform that connects genotypic, phenotypic, and environmental data for georeferenced plant populations. It is the primary platform of the lab and the exclusive destination for plant population data, and it grew directly out of TreeGenes.
From TreeGenes to CartograPlant: TreeGenes originated from the Dendrome Project in the mid-1990s as one of the first USDA Agricultural Research Service genome databases, built to centralize genetic resources for forest trees. As the platform expanded to include genomes, transcriptomes, population data, and environmental metadata, a gap emerged in the ability to integrate genotype, phenotype, and environment (GxPxE) in a spatial, population-level framework. CartograPlant was developed to address that need, providing a map-based interface for integrated analysis of georeferenced populations, traits, and genomic data. It maintains a focus on forest tree populations while housing GxPxE datasets not well represented elsewhere. Visiting treegenesdb.org now redirects to CartograPlant.
The idea began in June 2011, when forest tree biology researchers working in physiology, ecology, genomics, and systematics recognized the need for a unified platform to integrate and visualize spatial biological data. They came together through workshops and collaborations funded by the NSF iPlant Collaborative, now CyVerse, with the goal of bridging gaps between disciplines and making georeferenced population data accessible alongside traits and genotypes. The first version, then called CartograTree, was released in 2012 on iPlant infrastructure. By 2015 a more refined version allowed users to identify, filter, compare, and visualize spatial data across species distributions, genetic information, and environmental factors.
CartograPlant links over 20 million georeferenced records from more than 800 plant species with nearly 1,000 environmental layers, supporting both meta-analysis of published studies and joint analysis of raw datasets. It identifies genotype and environment associations, genotype and trait associations, population structure, and diversity estimates. The application is built on the Tripal framework, which integrates the Chado schema with Drupal-based content management and supports interoperability across organismal databases. Within that framework we develop and maintain custom Tripal modules and integrate workflows for metadata annotation, quality control, and analysis within a FAIR data ecosystem. Nextflow integration enables containerized, reproducible analyses across high-performance computing environments, including workflows for variant standardization across genome versions and for population structure, genetic diversity, and association genetics. API connections reach external trait resources such as the Botanical Information and Ecology Network (BIEN), and the platform ingests trait data collected in the field at both the individual accession level and across established monitoring plots, including observations contributed through TreeSnap, a citizen science application developed by scientists at the University of Kentucky and the University of Tennessee. The application is part of the USDA AgBioData consortium and serves over 2,700 active users, and our development team runs virtual and in-person workshops at least twice annually.