Transparent checkpoint and live migration middleware for
This technology provides a way for large computing clusters to save their work-in-progress ('checkpointing') without interrupting the programs running on them, and to move those programs between machines seamlessly. A lightweight software library sits on each machine and handles the checkpointing behind the scenes, invisible to the operating system and to other machines in the cluster. It uses a windowed message-logging approach to coordinate saves across many simultaneous processes efficiently. When a machine needs to be taken offline or a workload rebalanced, the library remaps hardware addresses so processes can resume on a different machine as if nothing happened.
What you could build
A fault-tolerance middleware layer sold or licensed to HPC cloud providers, financial institutions running distributed simulations, or government labs running long-duration compute jobs — anywhere downtime or job loss carries real cost.
Who in Virginia should care
Northern Virginia's massive cloud and federal data center ecosystem — AWS, Microsoft Azure, and defense contractors like Leidos and SAIC — would have natural interest in fault-tolerant distributed compute middleware.
Readiness: Prototype likely
Concept — described but not yet demonstrated. Lab validated — supported by experimental results in the patent. Prototype likely — the text describes a built, working embodiment.
Readiness is inferred from the patent text, not from a lab visit.
The record
- Inventors
- Srinidhi Varadarajan, Joseph Ruscio
- Granted
- November 28, 2017
- Status
- Granted patent
- Patent number
- 9830095
Ready to talk?
Virginia Tech Intellectual Properties handles licensing for this technology.
Prosim summaries are generated from public patent text and are not legal advice.