Skip to content

Repository files navigation

ROSAI.java

Practical JDK25 FFI/Llama.CPP with ROSJavaLite message bus (https://github.com/meta-llama/llama3), 3.1 and 3.2 inference.

This project is the successor of [llama4j.java]llama2.java based on [llama4j.java]llama2.c by Andrej Karpathy and his excellent educational videos.

Features

  • llama.cpp callout via FFI
  • GPU support via CUA, cuBLAS, Vulkan
  • GGUF format parser
  • Llama 3 tokenizer based on minbpe
  • Llama 3 inference with Grouped-Query Attention
  • Support Llama 3.1 (ad-hoc RoPE scaling) and 3.2 (tie word embeddings)
  • Support F16, BF16 weights + Q8_0 and Q4_0 quantizations
  • Simple CLI with --chat and --instruct modes.
  • GraalVM's Native Image support (EA builds here)
  • Java Message bus a la ROSJava

Interactive --chat mode in action:

Setup

Download pure Q4_0 and (optionally) Q8_0 quantized .gguf files from:

The pure Q4_0 quantized models are recommended, except for the very small models (1B), please be gentle with huggingface.co servers:

# Llama 3.2 (3B)
curl -L -O https://huggingface.co/mukel/Llama-3.2-3B-Instruct-GGUF/resolve/main/Llama-3.2-3B-Instruct-Q4_0.gguf

# Llama 3.2 (1B)
curl -L -O https://huggingface.co/mukel/Llama-3.2-1B-Instruct-GGUF/resolve/main/Llama-3.2-1B-Instruct-Q8_0.gguf

# Llama 3.1 (8B)
curl -L -O https://huggingface.co/mukel/Meta-Llama-3.1-8B-Instruct-GGUF/resolve/main/Meta-Llama-3.1-8B-Instruct-Q4_0.gguf

# Llama 3 (8B)
curl -L -O https://huggingface.co/mukel/Meta-Llama-3-8B-Instruct-GGUF/resolve/main/Meta-Llama-3-8B-Instruct-Q4_0.gguf

# Optionally download the Q8_0 quantized models
# curl -L -O https://huggingface.co/mukel/Meta-Llama-3-8B-Instruct-GGUF/resolve/main/Meta-Llama-3-8B-Instruct-Q8_0.gguf
# curl -L -O https://huggingface.co/mukel/Meta-Llama-3.1-8B-Instruct-GGUF/resolve/main/Meta-Llama-3.1-8B-Instruct-Q8_0.gguf

Optional: quantize to pure Q4_0 manually

In the wild, Q8_0 quantizations are fine, but Q4_0 quantizations are rarely pure e.g. the token_embd.weights/output.weights tensor are quantized with Q6_K, instead of Q4_0.
A pure Q4_0 quantization can be generated from a high precision (F32, F16, BFLOAT16) .gguf source with the llama-quantize utility from llama.cpp as follows:

./llama-quantize --pure ./Meta-Llama-3-8B-Instruct-F32.gguf ./Meta-Llama-3-8B-Instruct-Q4_0.gguf Q4_0

Build and run

Java 21+ is required, in particular the MemorySegment mmap-ing feature.

Run from source

java --enable-preview --source 21 --add-modules jdk.incubator.vector org.ros.ROSCore <ROS node class> -i --model Meta-Llama-3-8B-Instruct-Q4_0.gguf

Optional: Makefile + manually build and run

llama.cpp

Vanilla llama.cpp built with make.

./llama-cli --version                                                                                                                                                                          130 ↵
version: 3862 (3f1ae2e3)
built with cc (GCC) 14.2.1 20240805 for x86_64-pc-linux-gnu

Executed as follows:

./llama-bench -m Llama-3.2-1B-Instruct-Q4_0.gguf -p 0 -n 128

Llama3.java

taskset -c 0-15 ./llama3 \
  --model ./Llama-3-1B-Instruct-Q4_0.gguf \
  --max-tokens 128 \
  --seed 42 \
  --stream false \
  --prompt "Why is the sky blue?"

Hardware specs: 2019 AMD Ryzen 3950X 16C/32T 64GB (3800) Linux 6.6.47.

**Notes
Running on a single CCD e.g. taskset -c 0-15 ./llama3 ... since inference is constrained by memory bandwidth.

Results

License

MIT

About

JDK 25 FFI to Llama.cpp to RosJavaLite model runner for AI embodied robotics.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages