A prototype of Google’s orbit calculation satellite took off today aboard a SpaceX rocket from California. This is the first time a tech giant has sent its advanced chips into space.
The satellite, built by Planet Labs, proves that Nvidia’s GPU competitor, the Google Tensor Processing Unit, can function in space. That means you have to continuously provide 1 kilowatt of power, keep the chip cool, and run a series of models at a given pace to see if anything goes wrong.
“We’ve done tests on the ground, but no test is quite as good as the real thing,” said Travis Beals, an executive manager at Google who manages Project Suncatcher, the tech giant’s plan to develop large computing clusters in Earth orbit.
Once operational, the satellite will power up the TPU in 15-minute bursts to avoid taxing the satellite’s power and thermal management systems. The satellite is based on a standard platform built by Planet Labs, but the companies are working on a demo scheduled to fly next year that will see two satellites built specifically for advanced computing capable of running larger workloads. Future versions of these will attempt to work together via laser communication links.
Suncatcher isn’t the only space AI payload on board this SpaceX rocket, with over 100 different payloads launched, including missions from Satlyt and Cowboy Space Company.
What differentiates Google’s effort from those startups (and indeed SpaceX itself) is that it’s a long-term project.
The focus of this “long-term moonshot,” as Beals puts it, is on building the space infrastructure and AI workloads that will exist in the future. The company envisions an orbital data center, a network of 81 satellites flying in close formation and processing in parallel.
“Bandwidth and latency between TPUs is very important when you’re trying to run multi-rack workloads. We’re trying to look at not just what workloads exist today, but what workloads will be in five years,” said Beals. The main reason for this is that the rockets needed to cost-effectively scale up orbital data centers don’t yet exist.
Google also released a peer-reviewed version of its orbital data center white paper on Thursday. This is one of the most rigorous analyzes available of how computing reaches its trajectory. This paper will be published in the journal Joule.
One of the most interesting aspects of this paper is how Google thinks about access to space. Although the researchers stress that their analysis is not an economic feasibility study, the analysis provides an interesting picture of how the company sees rockets becoming cheaper over time.
Like other data center companies, Google is counting on SpaceX to launch its spacecraft. (Google is also a major investor in SpaceX.)
Elon Musk’s rocket maker claims to have achieved a “learning curve” of roughly 20% annual price reductions since the launch of its Falcon 1 rocket, and the authors think it is reasonable to expect the company to achieve launch prices close to $200 per kilogram by 2035.
What do you need for that? Based on the amount of payload launched by the Falcon 9, they believe Starship would need to fly 370,000 tons of payload into orbit to achieve a similar cost-saving trajectory. This would require about 1,800 launches over the next 10 years, or 180 per year, if each mission could fly 200 tons.
This is a big ask for a vehicle that has never flown more than five times in a year. SpaceX predicts the company will fly much more than that – Elon Musk, for example, has suggested Starship could reach flight speeds per hour in 2029, but Musk has said as much.
At least the good news from Google’s latest research is that its chips are likely to survive cosmic radiation. The company had to redo tests blasting the chip in a particle accelerator after realizing that the chip’s configuration provided more shielding than it experienced in real life. This introduced slightly more errors in the chip’s logic, but the company is still confident that its chip can handle large inference workloads in orbit for the five years of the satellite’s lifespan.
“If you think about common inference operations, the error rate is very low, about one in a million,” Beals says. “On the other hand, there were already problems when running large-scale training, for example, running thousands of chips for months.”
Correction: The headline for this article originally incorrectly stated the estimated learning curve for Starship launches as 1,600. 1,800.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
