Transparent checkpoint and migration middleware for HPC jobs
This technology allows a distributed computing system—where many computers work together on a single task—to periodically save the state of all running processes in a coordinated way, without requiring changes to the operating system. It uses a smart memory-tracking approach where each computer records which parts of memory have changed, saves those changes efficiently, and protects already-saved data from being overwritten before it is safely stored. If a computer node fails or needs to be taken offline, the saved checkpoint allows work to resume from a known good state, and processes can even be moved from one machine to another seamlessly. The entire mechanism operates invisibly at the software library level, making it relatively easy to deploy on existing systems.
What you could build
A fault-tolerance middleware library for high-performance computing (HPC) clusters, scientific computing environments, or large-scale cloud workloads that automatically checkpoints and migrates jobs; buyers would be HPC cluster operators, national labs, and cloud infrastructure providers needing job resilience without OS-level modifications.
Who in Virginia should care
Northern Virginia's massive data center corridor and federal HPC users (DOD, intelligence community contractors, AWS GovCloud) would have direct interest in fault-tolerant distributed computing middleware.
Readiness: Lab validated
Concept — described but not yet demonstrated. Lab validated — supported by experimental results in the patent. Prototype likely — the text describes a built, working embodiment.
Readiness is inferred from the patent text, not from a lab visit.
The record
- Inventors
- Srinidhi Varadarajan, Joseph Ruscio
- Granted
- May 19, 2009
- Status
- Granted patent
- Patent number
- 7536591
Ready to talk?
Virginia Tech Intellectual Properties handles licensing for this technology.
Prosim summaries are generated from public patent text and are not legal advice.