Data lakehouse Onehouse nabs $35M to capitalize on GenAI revolution

Paul Sawers

Updated 26 June 2024 at 6:17 pm·7-min read

You can barely go an hour these days without reading about generative AI. While we are still in the embryonic phase of what some have dubbed the "steam engine" of the fourth industrial revolution, there's little doubt that "GenAI" is shaping up to transform just about every industry — from finance and healthcare to law and beyond.

Cool user-facing applications might attract most of the fanfare, but the companies powering this revolution are currently benefiting the most. Just this month, chipmaker Nvidia briefly became the world's most valuable company, a $3.3 trillion juggernaut driven substantively by the demand for AI computing power.

But in addition to GPUs (graphics processing units), businesses also need infrastructure to manage the flow of data — for storing, processing, training, analyzing and, ultimately, unlocking the full potential of AI.

One company looking to capitalize on this is Onehouse, a three-year-old Californian startup founded by Vinoth Chandar, who created the open source Apache Hudi project while serving as a data architect at Uber. Hudi brings the benefits of data warehouses to data lakes, creating what has become known as a "data lakehouse," enabling support for actions like indexing and performing real-time queries on large datasets, be that structured, unstructured or semi-structured data.

For example, an e-commerce company that continuously collects customer data spanning orders, feedback and related digital interactions will need a system to ingest all that data and ensure it's kept up-to-date, which might help it recommend products based on a user's activity. Hudi enables data to be ingested from various sources with minimal latency, with support for deleting, updating and inserting ("upsert"), which is vital for such real-time data use cases.

Onehouse builds on this with a fully managed data lakehouse that helps companies deploy Hudi. Or, as Chandar puts it, it "jumpstarts ingestion and data standardization into open data formats" that can be used with nearly all the major tools in the data science, AI and machine learning ecosystems.

"Onehouse abstracts away low-level data infrastructure build-out, helping AI companies focus on their models," Chandar told TechCrunch.

Today, Onehouse announced it has raised $35 million in a Series B round of funding as it brings two new products to market to improve Hudi's performance and reduce cloud storage and processing costs.

Down at the (data) lakehouse

Chandar created Hudi as an internal project within Uber back in 2016, and since the ride-hailing company donated the project to the Apache Foundation in 2019, Hudi has been adopted by the likes of Amazon, Disney and Walmart.

Chandar left Uber in 2019, and, after a brief stint at Confluent, founded Onehouse. The startup emerged out of stealth in 2022 with $8 million in seed funding, and followed that shortly after with a $25 million Series A round. Both rounds were co-led by Greylock Partners and Addition.

These VC firms have joined forces again for the Series B follow-up, though this time, David Sacks' Craft Ventures is leading the round.

"The data lakehouse is quickly becoming the standard architecture for organizations that want to centralize their data to power new services like real-time analytics, predictive ML and GenAI," Craft Ventures partner Michael Robinson said in a statement.

For context, data warehouses and data lakes are similar in the way they serve as a central repository for pooling data. But they do so in different ways: A data warehouse is ideal for processing and querying historical, structured data, whereas data lakes have emerged as a more flexible alternative for storing vast amounts of raw data in its original format, with support for multiple types of data and high-performance querying.

This makes data lakes ideal for AI and machine learning workloads, as it's cheaper to store pre-transformed raw data, and at the same time, have support for more complex queries because the data can be stored in its original form.

However, the trade-off is a whole new set of data management complexities, which risks worsening the data quality given the vast array of data types and formats. This is partly what Hudi sets out to solve by bringing some key features of data warehouses to data lakes, such as ACID transactions to support data integrity and reliability, as well as improving metadata management for more diverse datasets.

Because it is an open source project, any company can deploy Hudi. A quick peek at the logos on Onehouse's website reveals some impressive users: AWS, Google, Tencent, Disney, Walmart, ByteDance, Uber and Huawei, to name a handful. But the fact that such big-name companies leverage Hudi internally is indicative of the effort and resources required to build it as part of an on-premises data lakehouse setup.

"While Hudi provides rich functionality to ingest, manage and transform data, companies still have to integrate about half-a-dozen open source tools to achieve their goals of a production-quality data lakehouse," Chandar said.

This is why Onehouse offers a fully managed, cloud-native platform that ingests, transforms and optimizes the data in a fraction of the time.

"Users can get an open data lakehouse up-and-running in under an hour, with broad interoperability with all major cloud-native services, warehouses and data lake engines," Chandar said.

The company was coy about naming its commercial customers, aside from the couple listed in case studies, such as Indian unicorn Apna.

"As a young company, we don’t share the entire list of commercial customers of Onehouse publicly at this time," Chandar said.

With a fresh $35 million in the bank, Onehouse is now expanding its platform with a free tool called Onehouse LakeView, which provides observability into lakehouse functionality for insights on table stats, trends, file sizes, timeline history and more. This builds on existing observability metrics provided by the core Hudi project, giving extra context on workloads.

"Without LakeView, users need to spend a lot of time interpreting metrics and deeply understand the entire stack to root-cause performance issues or inefficiencies in the pipeline configuration," Chandar said. "LakeView automates this and provides email alerts on good or bad trends, flagging data management needs to improve query performance."

Additionally, Onehouse is also debuting a new product called Table Optimizer, a managed cloud service that optimizes existing tables to expedite data ingestion and transformation.

'Open and interoperable'

There's no ignoring the myriad other big-name players in the space. The likes of Databricks and Snowflake are increasingly embracing the lakehouse paradigm: Earlier this month, Databricks reportedly doled out $1 billion to acquire a company called Tabular, with a view toward creating a common lakehouse standard.

Onehouse has entered a hot space for sure, but it's hoping that its focus on an "open and interoperable" system that makes it easier to avoid vendor lock-in will help it stand the test of time. It is essentially promising the ability to make a single copy of data universally accessible from just about anywhere, including Databricks, Snowflake, Cloudera and AWS native services, without having to build separate data silos on each.

As with Nvidia in the GPU realm, there's no ignoring the opportunities that await any company in the data management space. Data is the cornerstone of AI development, and not having enough good quality data is a major reason why many AI projects fail. But even when the data is there in bucketloads, companies still need the infrastructure to ingest, transform and standardize to make it useful. That bodes well for Onehouse and its ilk.

"From a data management and processing side, I believe that quality data delivered by a solid data infrastructure foundation is going to play a crucial role in getting these AI projects into real-world production use cases — to avoid garbage-in/garbage-out data problems," Chandar said. "We are beginning to see such demand in data lakehouse users, as they struggle to scale data processing and query needs for building these newer AI applications on enterprise scale data."

The Independent
I was at the Trump-Biden presidential debate and it became very clear what had gone wrong
When the debate ended, the Biden surrogates were nowhere to be found — until they emerged, grim-faced, as a group, and the rumors began. Andrew Feinberg reports on what happened behind the scenes at the first presidential debate in Atlanta, Georgia — and why the president’s performance went so badly
The Guardian
Who could replace Joe Biden? Here are six possibilities
With Biden not yet officially endorsed as Democratic presidential candidate, it is in theory open to the party to choose another candidate
Twentytwo13
Malaysians on tenterhooks after Johor Regent takes on Selangor Sultan in public domain
The football drama in Malaysia has reached fever-pitch, with the spotlight now on the Johor and Selangor palaces. The post Malaysians on tenterhooks after Johor Regent takes on Selangor Sultan in public domain appeared first on Twentytwo13.
The Independent
‘Panic mode’ Democrats begin calling for Biden to step aside after ‘horrible’ debate performance against Trump
‘Need to have Harris take over. Cleanest option,’ one Democrat strategist told The Independent
Evening Standard
Police investigate video 'showing female prison officer having sex with inmate in Wandsworth Prison cell'
The Ministry of Justice called in police after the video was released online
Futurism
China Finds Something Strange in Sample Retrieved From Moon
Lunar Graphene Chinese scientists have made an unusual discovery while analyzing the sample Chang'e-5 collected from the Moon's surface in December 2020. They found naturally occurring "few-layer graphene" for the first time, as state-run news agency Global Times reports, which could have major implications for our plans to make use of local resources once on […]
The Telegraph
‘This wasn’t a debate, it was a medical emergency’: Our writers give their verdicts
Follow the latest reaction in our live blog
Cosmopolitan
The Advice Matt Damon Reportedly Gave Ben Affleck as Things "Started Falling Apart" With J.Lo
Here's the advice Matt Damon reportedly gave Ben Affleck when things started "falling apart" with J.Lo.
People
Griff Is Giving Away the Dress She Wore to Open for Taylor Swift: ‘Wanted to Pass It Down’
The singer revealed she is giving away her "But Daddy I Love Him"-inspired dress she wore when she opened at the Eras Tour on June 22
CNN
China expels two former defense ministers from Communist Party as military purge deepens
China on Thursday expelled its former Defense Minister Li Shangfu from the ruling Communist Party over corruption allegations, state broadcaster CCTV reported, eight months after he was dramatically removed from the post.
SETHLUI.COM
Owners of private diner with 2 years waitlist open curry mee stall at Havelock Food Centre
The post Owners of private diner with 2 years waitlist open curry mee stall at Havelock Food Centre appeared first on SETHLUI.com.
INSIDER
US stockpiles of the rare earth minerals it would need to fight a war against an adversary like China are a mystery, and experts warn it's a problem
Whether the US rare earth minerals stockpile is ready for an emergency is shrouded in mystery, but there are indications it isn't.
INSIDER
A US Marine Corps attack helicopter fired off a new 'fire and forget' missile for the first time in the Pacific, striking a moving vessel
The precision "fire and forget" AGM-179 Joint Air-to-Ground Munition was fired at a moving sea target in waters off Japan.
Cosmo
Camila Cabello's shredded dress and tiiiiiny hot pants are Miami girl vibes
Camila Cabello is Complex's latest cover star. The artist models a see-through shredded dress as well as the tiniest hot pants in the accompanying photoshoot.
BBC
China honours woman who died saving Japanese family
Hu Youping, a bus attendant, tried to restrain an assailant at a bus stop outside a Japanese school.
INSIDER
A grinding Russian assault appears telling about Putin's plan to defeat Ukraine
ISW's conflict experts warned that the West must "challenge Putin's belief that he can gradually subsume Ukraine."
BBC
Dying together: Why a happily married couple decided to stop living
Jan and Els sought medical intervention to end their lives after 50 happy years of marriage.
HuffPost
10-Year-Old Headbanger Has Simon Cowell Saying 'Whoa!' In 'AGT' Audition
Maya Neelakantan strummed her guitar at first, then wailed away into the hearts of viewers.
BuzzFeed
Flight Attendants Are Sharing The Most Entitled Passengers They've Ever Had To Deal With, And I'm Disgusted
"Every so often, we get the odd straggler who boards last and finds a vacant seat in first or business, thinking that we won't know that they are from coach."
INSIDER
Singapore won best first-class airline in the world for its exclusive hotel-like Airbus A380 suite. Here's what it's like inside.
Singapore beat out Air France and Emirates for the world's best first-class, complete with a double bed. Tickets can cost up to $30,000 roundtrip.

Straits Times Index

Nikkei

Hang Seng

FTSE 100

Bitcoin USD

CMC Crypto 200

S&P 500

Dow

Nasdaq

Gold

Crude Oil

10-Yr Bond

FTSE Bursa Malaysia

Jakarta Composite Index

PSE Index

Data lakehouse Onehouse nabs $35M to capitalize on GenAI revolution

Down at the (data) lakehouse

'Open and interoperable'

Latest stories

I was at the Trump-Biden presidential debate and it became very clear what had gone wrong

Who could replace Joe Biden? Here are six possibilities

Malaysians on tenterhooks after Johor Regent takes on Selangor Sultan in public domain

‘Panic mode’ Democrats begin calling for Biden to step aside after ‘horrible’ debate performance against Trump

Police investigate video 'showing female prison officer having sex with inmate in Wandsworth Prison cell'

China Finds Something Strange in Sample Retrieved From Moon

‘This wasn’t a debate, it was a medical emergency’: Our writers give their verdicts

The Advice Matt Damon Reportedly Gave Ben Affleck as Things "Started Falling Apart" With J.Lo

Griff Is Giving Away the Dress She Wore to Open for Taylor Swift: ‘Wanted to Pass It Down’

China expels two former defense ministers from Communist Party as military purge deepens

Owners of private diner with 2 years waitlist open curry mee stall at Havelock Food Centre

US stockpiles of the rare earth minerals it would need to fight a war against an adversary like China are a mystery, and experts warn it's a problem

A US Marine Corps attack helicopter fired off a new 'fire and forget' missile for the first time in the Pacific, striking a moving vessel

Camila Cabello's shredded dress and tiiiiiny hot pants are Miami girl vibes

China honours woman who died saving Japanese family

A grinding Russian assault appears telling about Putin's plan to defeat Ukraine

Dying together: Why a happily married couple decided to stop living

10-Year-Old Headbanger Has Simon Cowell Saying 'Whoa!' In 'AGT' Audition

Flight Attendants Are Sharing The Most Entitled Passengers They've Ever Had To Deal With, And I'm Disgusted

Singapore won best first-class airline in the world for its exclusive hotel-like Airbus A380 suite. Here's what it's like inside.