# VGPU – Virtual GPU

**Computing as ODE Integration**  
*Author: Per Lindholm*  
*License: MIT*

---

## Overview

**VGPU** is a proof‑of‑concept framework that demonstrates a radical idea:

> **A GPU is an ODE solver running in parallel.**  
> **CUDA is a domain‑specific language for specifying many ODE trajectories at once.**

VGPU emulates CUDA kernels using pure ODE integration (RK4 / Euler), and provides a persistent kernel cache that compiles GLSL compute shaders once and reuses them for many dispatches.  
It bridges the gap between high‑level “field” specifications and real GPU hardware.

This repository contains:

- `ref3.py` – The original pure‑Python ODE integrator (CPU reference), which verifies equivalence to CUDA for scale, matmul, parallel trajectories, and sync.
- `vgpu-cache-full.py` – A persistent GLSL kernel cache that compiles a matmul shader once per shape and reuses it across many calls, with automatic CPU fallback.
- `vgpu-specification3.md` – The full theoretical specification of VGPU (version 1.2).

---

## The Core Thesis

Every CUDA primitive has a direct equivalent in the ODE world:

| CUDA | VGPU Equivalent |
|------|-----------------|
| `__global__ void kernel(...)` | `field { d*/dt = ... }` |
| `threadIdx.x` | Trajectory index |
| `blockIdx.x` | Basin seed |
| `__shared__` | Coupled sub‑field |
| `__syncthreads()` | Basin collapse threshold |
| `atomicAdd` | Flux‑conservative boundary op |

The implementation proves that matrix multiplication (GEMM) can be expressed as a set of independent accumulator trajectories, and that the result converges to the exact CUDA output.

---

## Files

| File | Description |
|------|-------------|
| `ref3.py` | Pure‑Python ODE interpreter with RK4/Euler; runs on CPU and verifies equivalence against CUDA references for four demos. |
| `vgpu-cache-full.py` | GPU‑accelerated matmul kernel cache using OpenGL compute shaders. Compiles once per `(M,K,N)` shape, reuses buffers and shader program. |
| `vgpu-specification3.md` | Complete theory, migration notes, and proof of CUDA equivalence (v1.2). |
| `c00.py` (optional) | Example integration of `VGPUCache` into a small MLP training loop. |

---

## Requirements

- Python 3.7+
- `numpy`
- `scipy` (only for the MLP example, not needed for the core demos)
- `moderngl` (for GPU acceleration; optional – falls back to CPU if unavailable)
- OpenGL 4.3+ with compute shader support (for GPU mode)

Install the mandatory packages:

```bash
pip install numpy moderngl
```

(For the MLP example, also `pip install scipy`.)

---

## Usage

### 1. CPU Reference (Equivalence Tests)

```bash
python ref3.py
```

This runs four tests:
- Scale kernel (`x[i] *= scale`)
- Matrix multiplication (`C = A @ B`, pure ODE accumulator)
- Parallel trajectories (convergence to `sin(h0)`)
- Barrier (`__syncthreads` basin collapse)

All tests should pass, proving that the ODE formulation matches CUDA exactly.

### 2. GPU‑Accelerated Matmul Cache

```bash
python vgpu-cache-full.py
```

This compiles the matmul shader for `N=16`, runs 100 calls (cache hits), then compiles for `N=64`, and finally interleaves calls to verify buffer rebinding.  
Expected output:

```
VGPUCache: 2 compiles, 198 cache hits, 200 GPU dispatches, 0 CPU fallbacks
```

### 3. Using `VGPUCache` in Your Own Code

```python
from vgpu_cache import VGPUCache   # from the fixed version

cache = VGPUCache()
C = cache.matmul(A, B)   # A: (M,K), B: (K,N) -> C: (M,N)
```

It automatically chooses GPU if available, otherwise falls back to `numpy @`.

---

## How It Works

### CPU Reference (`ref3.py`)

The `VGPU` class integrates a user‑defined vector field `F(h, t, h0)` using either RK4 or Euler.  
Each trajectory is independent; the class launches `N` trajectories in a loop (simulating parallel threads).

For matrix multiplication, each `C[i,j]` is a separate trajectory that accumulates `Σ_k A[i,k]*B[k,j]` over time.  
The field is piecewise‑constant, so Euler integration gives exact results.

### GPU Cache (`vgpu-cache-full.py`)

- **Compile once**: The GLSL compute shader is compiled the first time a shape `(M,K,N)` is seen.
- **Bind once**: Buffers for A, B, C are allocated per shape.
- **Rebind before dispatch**: Because OpenGL SSBO bindings are global state, we rebind the correct buffers before each `run()` call. This avoids cross‑kernel pollution.
- **Cache hits**: Subsequent calls with the same shape reuse the compiled program and buffers; only new data is uploaded (DMA) and the kernel is dispatched.

The cache is lazy: kernels are built only when first needed.

---

## License

This project is licensed under the **MIT License** – see the [LICENSE](LICENSE) file for details.

---

## Author

**Per Lindholm**  
*Contact*: [per.lindholm@example.com] (replace with actual email if desired)

---

## Acknowledgments

This work is inspired by the XYFLOW / ODE‑CCT framework and the idea that all computation can be seen as integrating trajectories through a vector field.

---

## Future Directions

- Extend the GPU cache to support arbitrary element‑wise operations (ReLU, softmax, etc.) via fused kernels.
- Add support for shared memory / cooperative groups.
- Build a full PTX‑to‑field compiler using the equivalence theorem.

---

*Last updated: 2026*
