Case Study


About Snowflake
Snowflake is the platform for the AI era, making it easy for enterprises to innovate faster and get more value from data. More than 13,900 customers around the globe, including hundreds of the world's largest companies, use Snowflake's AI Data Cloud to build, use and share data, applications and AI. With Snowflake, data and AI are transformative for everyone. Learn more at snowflake.com (NYSE: SNOW).
~50%
4 of 5
The operational overhead of self-managed RBE at scale
More than two years ago, Snowflake's Developer Productivity team migrated Snowflake to Bazel and RBE, running internally managed clusters across large Java and C++ codebases. The team involved comprised experienced Bazel engineers, including former members of the Google Bazel team.
At that stage, Snowflake had what many engineering organizations aspire to: modern build infrastructure, deep internal expertise, and full control over the platform.
As Snowflake scaled to over 2,000 engineers and a multi-million-line codebase, however, the operational requirements of maintaining and tuning their self-managed RBE toolchain grew along with the organization.
This meant allocating internal engineers to continuously manage cloud resources to ensure predictability, perform complex platform upgrades, and scale infrastructure as demand rose.
This was an operations challenge, not a talent problem.
Snowflake did what strong engineering teams do. They tuned the system, applied more care to upgrades, and focused efforts on stability, cost consciousness, and performance.
But the effort came with a considerable opportunity cost: in order to keep moving fast, Snowflake needed to allocate up to 5 elite developer productivity specialists at a time focusing on infrastructure maintenance.
So, Snowflake's leadership reframed the decision: the question was no longer whether the team could run RBE itself, as it clearly could. It was whether operating a high-cost distributed system on the company's critical path was the best use of those engineers.
Migrating to EngFlow's managed service
Build execution at this level is its own operational discipline, and Snowflake treated it that way. Snowflake signed with EngFlow in July, 2025 and had moved the bulk of its RBE workloads to EngFlow by October, 2025.
EngFlow changed the equation in two ways.
First, Snowflake moved to a managed service model, retaining the benefits of Bazel and remote execution while EngFlow absorbed the operational complexity of running the platform.
Second, EngFlow brought technical capabilities that directly addressed Snowflake's pain points.
Managing remote caching at Snowflake's scale is highly resource-intensive. EngFlow's remote execution and cache architecture, featuring automatic spillover to Amazon S3, allowed Snowflake to meet optimal time-to-live (TTL) windows and reduce storage costs.
EngFlow also made previously difficult platform capabilities straightforward. For example, it had been impractical to offer engineering teams a choice of execution base images because each additional image expanded the platform's operational surface. With EngFlow, adding an image became a mere configuration change.
The impact: almost 50% lower AWS costs and optimally allocated engineers
Snowflake now runs 17 million builds and 290 million actions per week on EngFlow.
EngFlow's autoscaling, scheduling, cache architecture, and data-transfer efficiency led to a nearly 50% reduction in AWS costs for cloud infrastructure compared with the prior self-managed setup, while build performance improved.
The staffing model changed just as dramatically.
Before EngFlow, Snowflake needed as many as five engineers at a time supporting RBE. After the migration, the number dropped to roughly half an engineer, with EngFlow managing the operational complexity.
This new stability also unlocked progress. Onboarding new workloads stopped being an event that could delay production deployments.
With the operational overhead now greatly reduced by comparison, Snowflake's engineering talent was able to refocus on developer productivity initiatives; for example, by prioritizing stricter development sandboxing, the team was able to expose build failures due to flaky tests that had been particularly hard to identify and diagnose.
Matching the cost of open source with the value of managed build infrastructure
Choosing a managed production service for infrastructure has become essential to Snowflake's software development.
Yet it's clear that Snowflake's story is not about open source being inadequate. It's about what happens when a critical internal platform becomes too important, too expensive, and too operationally complex to manage as a side responsibility.
As mature organizations will know, the cost of self-managing remote execution is not just cloud spend (though left unchecked at scale, this becomes an unpredictable expense). The opportunity cost comes from allocating experienced engineering talent toward infrastructure management rather than developer productivity and business innovation.
By moving to EngFlow, Snowflake freed its Bazel experts to focus on strategic developer productivity initiatives, stabilized a mission-critical system, helped reduce AWS costs substantially, and restored confidence in remote build execution.
