Investing in Synthefy: Building Foundation Models for the World's Numbers
Every platform shift in AI so far has been narrated in language. We prompt models in words, we read their answers in words, and the public imagination of AI runs almost entirely on text and speech. All the while the largest data modality on earth sits mostly untouched. Numbers, the raw output of machines, meters, and markets, pour out without pause while quietly driving how the physical and financial economy operates.
The scale here is hard to overstate. Google now processes over 3.2 quadrillion tokens a month across its surfaces, up from roughly 480 trillion a year earlier and 9.7 trillion two years before that. Numerical observations from machines and markets dwarf even that language-scale load. Each of those observations is to a number token what a word is to a language token, and they accumulate every second of every day, because the machines and markets producing them never idle.
Each of those numbers is a decision waiting to happen. Buried in that stream is a warning the business needs: equipment heading for failure, inventory about to run dry, a payment that should never clear. And the value sitting on that data is enormous and largely unrealized. Siemens estimates unplanned downtime will cost Fortune Global 500 industrial companies almost $1.5 trillion. Sharper forecasts on failure and related operational signals can claw much of that back, which hinges on data that current models simply cannot parse.
The category taking shape around that gap is tabular foundation models: pretrained models for structured and tabular data that learn to read tables the way large language models read text, and that can often predict on a new dataset without training a bespoke model for every table. Synthefy frames its work as Structured Data Foundation Models (SD FMs) inside that broader shift toward foundation models for the world's numbers.
This is the opportunity Synthefy was built for. We are proud to share that Wing led Synthefy's $6.5 million seed round to build SD FMs for structured data. Language already has its model and numbers are next.
Table of contents
- The team behind Synthefy
- Why tabular foundation models matter now
- From research to shipped models (Nori)
- Synthefy's opportunity
- Building the infrastructure layer for AI-driven numerical prediction
The team behind Synthefy: AI researchers from Uber, Stanford, and UT Austin
Synthefy is the kind of team that will win a category like this with an applied research group that has lived inside the numerical problem for years.
Somi Agarwal, Synthefy’s co-founder and CEO, worked on autonomous systems at Uber's ATG division and then went to UT Austin for a PhD focused on machine learning for time series. His research produced Time Weaver, an early paper on conditional time-series generation with heterogeneous metadata that anticipated the approach Synthefy now productizes at scale. That work sits next to the wider rise of foundation models for numerical modalities, from time series through tables.
Sandeep Chinchali, co-founder, is a professor at UT Austin and a Stanford CS PhD who advised Somi's doctoral work. He is actively contributing to Synthefy’s research, a signal of how convinced this group is that the research moment has arrived.
Raimi Shah, co-founder, brings systems depth from his work as a principal ML manager at Zscaler, with MS/BS training from UIUC.
In a category where research is the hard problem, the quality of the team's judgment is the product. This group has been asking how to teach a model to read numbers since before there was a market for the answer.
Why tabular foundation models matter now
Tables do not behave like language. Rows can be reordered without changing meaning. Schemas differ from company to company. The same digit can mean price, pressure, or probability depending on the column. Tricks that transfer cleanly across text break down when every dataset is a new schema and every cell needs the right relational context.
Enterprises still meet that mess with a familiar default. For each high-value table, teams train or tune a model—often gradient-boosted trees or a custom pipeline—then repeat the cold start when the schema, segment, or question changes. The cost is not only compute. It is feature work, validation, and maintenance multiplied across every new use case.
Tabular foundation models attack that loop at the root. They pretrain across large distributions of tables, often synthetic ones, so the model learns how to learn from structure rather than memorize a single spreadsheet. At inference, many of these systems use in-context learning: the labeled rows you provide act as the examples, and the model produces predictions on new rows without a full custom training program. The economic implication is the one we care about as infrastructure investors. Prediction on structured data starts to look less like a project and more like a shared capability—closer to an API call than a multi-quarter rebuild.
Research milestones such as TabPFN in Nature and industry efforts like Google’s TabFM show that the research moment is real. The open question is who turns that moment into a durable, builder-accessible layer for the numerical economy.
From research to shipped models: how Synthefy's Nori model outperforms existing ML pipelines
Synthefy did what the strongest research-led companies do. They published and shipped in the open, and the results speak for themselves.
Nori, their tabular foundation model for structured data, is the product expression of the category above: a shipped answer to “point a pretrained model at a table instead of standing up another pipeline.” It launched with open weights under Apache 2.0 and open training code. On public regression suites, the Hugging Face model card reports strong results across 96 tasks, and Synthefy’s product materials describe zero-shot wins against tuned XGBoost and LightGBM, with Nori-30M roughly 50× smaller than Google’s 1.6B TabFM. Because the model uses in-context learning, teams can point it at a new table without customer-specific training; hosted warm requests land on the order of seconds.
Two capabilities set this work apart. First, Synthefy’s models can make predictions on a new dataset out of the box. Point Nori at a table and it produces predictions in a single forward pass, without customer-specific feature engineering or model training. The work of building and tuning a bespoke ML pipeline becomes an API call.
Second, the models can combine structured data with unstructured text. In financial markets, they can read stock prices and trading data alongside earnings reports, regulatory filings, and news. In industrial operations, they can pair sensor readings with maintenance logs and incident reports. That multimodal edge matters because pure-tabular research models often stop at the grid, while real operations attach context in logs, filings, and tickets. This brings the context behind the numbers into each prediction and moves Synthefy toward one pretrained model that works across financial, commercial, and operational problems.
Synthefy's opportunity: teaching AI models to read numbers at scale
Most companies still turn structured data into predictions through a patchwork of task-specific systems. High-value problems get bespoke machine-learning models; many others remain in spreadsheets, business rules, or packaged software. Each new dataset and use case requires another round of setup, tuning, and maintenance. Tabular foundation models replace that repeated setup with one shared, pretrained model that reads tables the way an LLM reads language.
The pull is already visible. Teams in retail, finance, observability, and defense-oriented settings are exploring Synthefy for demand forecasting, pricing optimization, and failure prediction. Much of the interest arrives organically, from groups in fraud, retail, and industrial operations all asking the same question: can this model beat the one we built in-house on a use case that matters? Partnerships with Baseten and integrations with Snowflake and AWS SageMaker put the models into the platforms where that numerical data already lives. The largest technology companies are beginning to build tabular and structured-data foundation models too, which tells us the category is real. We believe the winning foundation for numbers will be open and independent, putting the model directly in builders’ hands so any team can point it at any numerical problem.
Building the infrastructure layer for AI-driven numerical prediction
Wing has invested in the data and AI infrastructure layer through every recent platform shift, and we recognize this shape. A research-led team with a data advantage that compounds, building a horizontal foundation and selling first to the most demanding buyers, tends to own the layer everyone else builds on once the category matures. We saw that pattern in Synthefy and we’re proud to lead this round.
We believe foundation models for structured and tabular data will become one of the defining infrastructure layers of the coming decade. It will sit quietly beneath demand plans, maintenance schedules, pricing engines, and fraud defenses at companies most people never think about, doing work that rarely makes headlines and matters enormously. Synthefy has the research depth, early traction, and conviction to build it. We’re proud to partner with Somi, Sandeep, Raimi, and the entire team as they teach machines to read the language of numbers. For more on how we think about modality shifts beyond language models, see From Words to Worlds and our other Insights.




