Back to the portfolio

Transparent checkpoint and live migration middleware for

Computing, Software & AIDefense & Aerospace ApplicationsPrototype likely

This technology provides a way for large computing clusters to save their work-in-progress ('checkpointing') without interrupting the programs running on them, and to move those programs between machines seamlessly. A lightweight software library sits on each machine and handles the checkpointing behind the scenes, invisible to the operating system and to other machines in the cluster. It uses a windowed message-logging approach to coordinate saves across many simultaneous processes efficiently. When a machine needs to be taken offline or a workload rebalanced, the library remaps hardware addresses so processes can resume on a different machine as if nothing happened.

What you could build

A fault-tolerance middleware layer sold or licensed to HPC cloud providers, financial institutions running distributed simulations, or government labs running long-duration compute jobs — anywhere downtime or job loss carries real cost.

Who in Virginia should care

Northern Virginia's massive cloud and federal data center ecosystem — AWS, Microsoft Azure, and defense contractors like Leidos and SAIC — would have natural interest in fault-tolerant distributed compute middleware.

Readiness: Prototype likely

Concept — described but not yet demonstrated. Lab validated — supported by experimental results in the patent. Prototype likely — the text describes a built, working embodiment.

Readiness is inferred from the patent text, not from a lab visit.

The record

Inventors
Srinidhi Varadarajan, Joseph Ruscio
Granted
November 28, 2017
Status
Granted patent
Patent number
9830095

Ready to talk?

Virginia Tech Intellectual Properties handles licensing for this technology.

VTIP contact coming shortly

Prosim summaries are generated from public patent text and are not legal advice.