Realtime LLM Chat on an 8GB Nvidia GPU
Realtime LLM Chat on an 8GB Nvidia GPU
Demo runs on a laptop 3070 Ti / 8GB. GPU memory doesn't go above 6GB, so it might run on an even smaller GPU. Uses a 4-bit 7bn parameter alpaca_lora model and performance is significantly worse than ChatGPT as you'd expect.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Incorrect prediction on native model
Similar products
trex_exporter – Prometheus Exporter for T-Rex Nvidia GPU Miner
Chisel – Profile GPU Kernels Without a GPU (Nvidia and AMD)
Nvidia Jetson Cameras and NVR
Can your GPU run this LLM?
Tilery-VM – run Nvidia cuTile GPU kernels on a CPU, no GPU required
Nvcachetools – See compiled shader code for Nvidia GPU, using the cache
Tunes – Rust audio synthesis/playback (100x realtime, SIMD, GPU, WASM)
Pebble realtime bus departures in Sydney
Realtime Blackboard
Realtime busses (websockets, leaflet and d3) (