Transparent checkpointing middleware for fault-resilient
This technology is a software system that creates periodic 'snapshots' of running processes across a cluster of computers, allowing work to be resumed from a saved point if something goes wrong — without requiring any changes to the operating system or applications. It uses a dual-buffer technique to capture memory changes in the background while computation continues uninterrupted, minimizing performance loss during the snapshot process. If a process or server fails, the system can roll back to the last checkpoint rather than restarting from scratch. The system can also move running processes from one machine to another by remapping network addresses, enabling load balancing or hardware maintenance without downtime.
What you could build
A fault-tolerance middleware layer for HPC clusters and cloud batch computing environments, sold to operators of scientific computing infrastructure, financial modeling farms, or large-scale simulation environments who need job resilience without rewriting applications.
Who in Virginia should care
Northern Virginia data center operators and federal HPC contractors (e.g., defense labs, intelligence community cloud programs) running large-scale distributed workloads would have natural interest.
Readiness: Concept
Concept — described but not yet demonstrated. Lab validated — supported by experimental results in the patent. Prototype likely — the text describes a built, working embodiment.
Readiness is inferred from the patent text, not from a lab visit.
The record
- Inventors
- Srinidhi Varadarajan, Joseph Ruscio
- Granted
- July 16, 2013
- Status
- Granted patent
- Patent number
- 8489921
Ready to talk?
Virginia Tech Intellectual Properties handles licensing for this technology.
Prosim summaries are generated from public patent text and are not legal advice.