Clockwork.io raises $31 million for AI cluster fault tolerance
Clockwork.io raised $31 million in a funding round to expand deployments of its fault-tolerance software for AI workloads. The company says it will extend deployments across training, inference and reinforcement learning, including through cloud partners.

Key takeaways · 4
- 01
Clockwork raised $31 million, bringing its total funding to $73 million.
- 02
Funding is intended to extend fault-tolerance deployments across training, inference and reinforcement learning, including through cloud partners.
- 03
TorchSnap snapshots distributed inference workloads across cluster nodes without code changes by developers.
- 04
LinkedIn has deployed Clockwork’s LinkPass across its GPU fleet; the feature reroutes AI traffic around optical-link and switch failures.
Funding and deployment plans
The $31 million round was co-led by Seligman Ventures, Wing Ventures and Premji Invest. Existing investors New Enterprise Associates and e& Capital also participated, bringing Clockwork’s total funding raised to $73 million.[1] Clockwork said it will use the funding to expand fault-tolerance deployments across training, inference and reinforcement learning, including through cloud partners.[2] SiliconANGLE dated its report on the round October 5, 2026.[1] The announcement describes Clockwork as a data-center infrastructure startup focused on improving the efficiency of AI chip clusters.[1]
Snapshots without code changes
Clockwork announced TorchSnap, a feature designed to minimize wasted compute.[1] It captures multinode snapshots of distributed AI inference workloads across cluster nodes without requiring developers to change their code.[1] Separately, fast asynchronous application checkpoints are aimed at reinforcement learning.[3] These are distinct details: the evidence describes TorchSnap in connection with distributed inference snapshots, and describes asynchronous checkpoints as aimed at reinforcement learning.[1][3] Clockwork’s broader fault-tolerance software is designed to keep training, reinforcement-learning and inference workloads running through infrastructure issues.[4]
Examples of deployment
LinkedIn has deployed Clockwork’s LinkPass functionality across its entire GPU fleet.[1] LinkPass reroutes AI traffic around optical-link and switch failures to keep jobs running.[1] Hiremagalur said the software prevents tens of thousands of GPU-hours of downtime each month across LinkedIn’s fleet.[3] WhiteFiber uses Clockwork technology to audit and validate cluster reliability before new AI workloads enter production.[1] These examples show different uses described in the evidence: routing around network-related failures and checking cluster reliability before production workloads are added.[1]
Why recovery time matters
Clockwork’s software layer sits between GPUs and running AI workloads, using nanosecond-accurate telemetry to identify failures before a full cluster restart.[1] Fault tolerance is intended to keep GPUs doing useful work rather than waiting for recovery or repeating completed work.[1] Clockwork said conventional recovery from a saved snapshot can take up to 90 minutes.[1] The evidence also notes that hardware failures can occur daily, and cites Meta’s report of hardware issues every three hours on average during Llama 3’s 54-day training run across 16,384 GPUs.[1]
For teams running AI workloads, the relevant question is whether fault tolerance can reduce time lost to infrastructure failures without requiring changes to application code. The funding and deployment examples make reliability a procurement and operations consideration, but the evidence does not establish results across customers or workloads beyond those described.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
6 October 2026
Clockwork.io raises $31 million for AI cluster fault tolerance
Sources
- A message from John Furrier, co-founder of SiliconANGLE:siliconangle.com
- Clockwork raises $31M as LinkedIn deploys LinkPass across its AI infrastructureruntimewire.com
- Clockwork Raises $31M as LinkedIn GPUs Ride Out Link Flapssupercomputing.news
- Clockwork.io Raises $31M as LinkedIn, Together AI and WhiteFiber Adopt Its Resilience Software to Stop Wasting GPU-Hoursprnewswire.com
- Clockwork Launches FleetIQ & Appoints Suresh Vasudevan as CEO - clockworkclockwork.io