Skip to main content

Clockwork.io raises $31 million for AI cluster fault tolerance

6 OCTOBER 2026·2 MIN READ·5 SOURCES

Clockwork.io raised $31 million in a funding round to expand deployments of its fault-tolerance software for AI workloads. The company says it will extend deployments across training, inference and reinforcement learning, including through cloud partners.

Clockwork.io raises $31 million for AI cluster fault tolerance

Key takeaways · 4

  • 01

    Clockwork raised $31 million, bringing its total funding to $73 million.

  • 02

    Funding is intended to extend fault-tolerance deployments across training, inference and reinforcement learning, including through cloud partners.

  • 03

    TorchSnap snapshots distributed inference workloads across cluster nodes without code changes by developers.

  • 04

    LinkedIn has deployed Clockwork’s LinkPass across its GPU fleet; the feature reroutes AI traffic around optical-link and switch failures.

Funding and deployment plans

The $31 million round was co-led by Seligman Ventures, Wing Ventures and Premji Invest. Existing investors New Enterprise Associates and e& Capital also participated, bringing Clockwork’s total funding raised to $73 million.[1] Clockwork said it will use the funding to expand fault-tolerance deployments across training, inference and reinforcement learning, including through cloud partners.[2] SiliconANGLE dated its report on the round October 5, 2026.[1] The announcement describes Clockwork as a data-center infrastructure startup focused on improving the efficiency of AI chip clusters.[1]

Snapshots without code changes

Clockwork announced TorchSnap, a feature designed to minimize wasted compute.[1] It captures multinode snapshots of distributed AI inference workloads across cluster nodes without requiring developers to change their code.[1] Separately, fast asynchronous application checkpoints are aimed at reinforcement learning.[3] These are distinct details: the evidence describes TorchSnap in connection with distributed inference snapshots, and describes asynchronous checkpoints as aimed at reinforcement learning.[1][3] Clockwork’s broader fault-tolerance software is designed to keep training, reinforcement-learning and inference workloads running through infrastructure issues.[4]

Examples of deployment

LinkedIn has deployed Clockwork’s LinkPass functionality across its entire GPU fleet.[1] LinkPass reroutes AI traffic around optical-link and switch failures to keep jobs running.[1] Hiremagalur said the software prevents tens of thousands of GPU-hours of downtime each month across LinkedIn’s fleet.[3] WhiteFiber uses Clockwork technology to audit and validate cluster reliability before new AI workloads enter production.[1] These examples show different uses described in the evidence: routing around network-related failures and checking cluster reliability before production workloads are added.[1]

Why recovery time matters

Clockwork’s software layer sits between GPUs and running AI workloads, using nanosecond-accurate telemetry to identify failures before a full cluster restart.[1] Fault tolerance is intended to keep GPUs doing useful work rather than waiting for recovery or repeating completed work.[1] Clockwork said conventional recovery from a saved snapshot can take up to 90 minutes.[1] The evidence also notes that hardware failures can occur daily, and cites Meta’s report of hardware issues every three hours on average during Llama 3’s 54-day training run across 16,384 GPUs.[1]

For teams running AI workloads, the relevant question is whether fault tolerance can reduce time lost to infrastructure failures without requiring changes to application code. The funding and deployment examples make reliability a procurement and operations consideration, but the evidence does not establish results across customers or workloads beyond those described.

Why it matters
Daily session

Put this to work — one session a day, built for your industry.

Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.

Start free

How this developed

  1. 6 October 2026

    Clockwork.io raises $31 million for AI cluster fault tolerance

Sources

AI fluency, one session a day, built for your work.