Building the first Large Geospatial Model to answer the most difficult questions about our planet.
- Contact
- contact@columbus.earth
- Social
Building a brain for Earth
At Columbus, we collect the world’s physical data, and build a model that comprehends it all.
We’re building frontier geospatial intelligence.
An Interactive Thesis
Large Geospatial Model vs Large Language Model.
Columbus’ primary research objective is to create a general intelligence for the physical world. We’re building foundation AI to understand the physical world, and the structure, semantics and relationships within it.
Where a language model learns the patterns and relationships in text, a Large Geospatial Model (LGM) learns the patterns and relationships in geographical spaces. Instead of language and text, we process data about our surroundings and the anthropology within them. We’re personifying Earth with a world-smart AI.
The physical world needs physical AI. We learned first-hand how LLM architecture was unable to consistently satisfy the queries and edge cases of our users. These limitations compelled us to build a generally intuitive model that could answer our users' hardest questions about our physical world.
Roughly 85% of global GDP activity happens in the physical world through logistics, construction, transportation, agriculture, resource extraction and more. However, the last few years of AI progress have focused on language, code, and images. The single largest domain of human activity is still waiting for its foundation models.
What is an LGM?

An LGM vs other foundation models
LLM
Large-Language-Model
VLM
Vision-Language-Model
LGM
Large-Geospatial-Model
Text
e.g. “The grass is” → “green”
Text & Image
e.g. dog photo → “a border collie”
Physical reality
e.g. public data + urban imagery + GIS → crime risk map
outputs
Predictive words
“What word comes next?”
Visual reasoning
“What’s in this image?”
Ground truths
“What’s in this physical space?”
building it







Columbus EarthA Large Geospatial Model is the next frontier in AI
The path forward
& Vision Models

LGM
Timeline of foundational AI models
- 2022 — LLM (Large Language Model)
- 2025 — Geo-tuned LLM & Vision Models
- 2027 — Generalist LGM (Large Geospatial Model)
- 2028 — UGM (Universal Geospatial Model)
- 2022LLM
- 2025Geo-tuned LLM & Vision Models
Now — August 2026- 2027Generalist LGM
- 2028UGMOur Game Plan
The Large Geospatial Model: Our Timeline

Our Architecture
We learned first-hand how Language Models consistently failed to satisfy geospatial queries and features. Even with perfect retrieval, we encountered issues that go beyond insufficient data:
- No native representation of distance, adjacency, or terrain. A latitude/longitude pair is processed as a sequence of sub-word tokens.
- No native geometric or topological awareness. 3D, interconnected space is serialized into a 1D linear string of text, so scaling the context window does not scale spatial intelligence.
- No authentic semantic core reasoning on geospatial data. No model reliably understands about physical location what’s intuitively obvious to us, especially in edge cases.
We were compelled to develop our proprietary architecture.
A system comprised of 3 parts: Data Collection, Fusion, and Core Reasoning. Within each are several novel approaches formulated through our applied research, including Data Enzymes and Earth Recipes.
Core Reasoning includes our concept of Earth Recipes: expert-agents that collaborate using our relational-vector architecture.
Why LLMs Didn't Cut It
In this interactive visualisation, we describe several details of our work so far. Tap for an overview of each layer:
We pride ourselves on the most extensive data collection in the industry. We find creative and versatile methods to source data, including from drone-imagery, car-cameras, public data, private data, user-crowdsourced data, and partner data providers.
Powered by our Data EnzymesData EnzymesJust as your stomach breaks down food into the nutrients your body can actually use, our data digestion pipeline takes raw, messy, dormant inputs and turns them into clean, structured, well-labeled data the model can learn from. technology, we’re able to expand our data collection to capture previously-untapped sources. Using our data digestion engine, we’re achieving the cheapest P/POI.P/POI — Price per Point of InterestThe cost to capture a single geospatial data point. Lower P/POI means richer, more affordable training data, a key economic metric for any geospatial foundation model. Data Enzymes allow us to turn dormant data usable, and to provide lower-cost enriched data sets to a wider customer base.
Read about our Data Digestion
Data Collection and Fusion are the most important steps in our architecture. That’s why we’re building automatic and accurate data filteringData filteringThe process of selecting, refining, and removing unwanted information from a dataset, such as noise, duplicates, or irrelevant variables. Additionally, scoring each data point for confidence and cross-validating it against other sources. and labelling.Automatic labellingThe process of tagging raw, unstructured geospatial data with structured metadata so a model can reason about it. Our pipeline does this without human annotation.
Data scarcity and data quality, the Pinnacle Data Problem,The Pinnacle Data ProblemThere’s not enough high-quality and well-labeled structured data currently available for a truly general-applicable Large Geospatial Model. Insufficient or lacking-in-quality training data is a pinnacle bottleneck to foundation models today. The data is out there but currently hard to capture. is one of the hardest parts about building the LGM. To solve this we have built distinct methods to universally digest data. With Data Enzymes,Data EnzymesJust as your stomach breaks down food into the nutrients your body can actually use, our data digestion pipeline takes raw, messy, dormant inputs and turns them into clean, structured, well-labeled data the model can learn from. we’re able to fuse data together for our model to train on and reason over. Cheaper to harvest → more data → smarter model.
We care about Ground Truths.Ground truthA verified, real-world observation confirmed at a specific X, Y, Z coordinate and time. We vet our data with accredited organizations, as well as conduct internal validation audits,Internal validation auditsMethodology for internal validation audits are always developing. Currently, we perform cross-validation against proxy variables as well as probabilistic consistency checks, to assess whether the model’s predictions are statistically meaningful. In addition, we routinely conduct proprietary ground surveying to validate data with certainty. to ensure each attribute is truthful at a given X,Y,Z point.
Read about our Data Enzymes
Our core reasoning architecture considers temporal data,Temporal dataObservations carrying a timestamp. For example, a street’s road condition in 2012 and the same street today are separate observations; keeping both allows a model to learn an attribute timeline, which helps it reason about how a place has been changing, and its future prediction. and sifts through vast amounts of aggregated geospatial data, including anthropologic data.
Our model includes our concept of Earth Recipes,Earth RecipesA recipe is a set of rules for how something comes to be. Copper deposits, for instance, form where particular geological conditions coincide. Rather than asking where its training data happened to label copper, the model asks what the recipe for copper is, which attributes it depends on, and where on Earth those attributes line up, inferring a new situation from rules and reality. which are fed to expert-agentsExpert-agentsExpert agents in AI are autonomous, specialized entities within the model, designed to execute complex tasks, execute decisions, and use external tools within a specific domain. They combine domain-specific Earth Recipes knowledge with goal-driven reasoning. They work together in parallel and combine to reach conclusions. that collaborate to learn and reason. MagellanMagellanMagellan is the name for the flagship Large Geospatial Models currently in development by Columbus Earth. continuously learns and creates new patterns through a relational architecture.
Read about our Earth Recipes
One model, innumerable applications.
Data Collection
The most extensive data collection in the industry. Versatile methods ranging from drones, car data, human data, public data and more.
Read our blogEarth Recipes: How a World Would Think
Our Model: Magellan-1.0
Capabilities in development
Contextually enriched reasoning over the semantics of urban space
Generative geospatial data
A generalist model, grounded on a living data catalogue
Granular reasoning at scale
Read our latest releases
Explore the innovative research and recent papers from our team.

Earth Recipes: How a World Would Think
Building accurate geospatial core reasoning with our concept of Earth Recipes

What is a Large Geospatial Model?
Definition of a Large Geospatial Model, and how it differs from World Models, LLMs and other foundation models

The Philosophy of a Universal Geospatial Model
Our vision for a universally capable model for spatial and geospatial reasoning
Careers
If you're excited about creating paradigm shifts in physical world understanding.
Join the crew
Research freely at Columbus.