Skip to content

neocoretechs/ROSAI

Repository files navigation

ROSAI.java

Practical JDK25 FFI/Llama.CPP with ROSJavaLite message bus (https://github.com/meta-llama/llama3), 3.1 and 3.2 inference.

This project is the successor of [llama4j.java]llama2.java based on [llama4j.java]llama2.c by Andrej Karpathy and his excellent educational videos.

Features

  • llama.cpp callout via FFI
  • GPU support via CUA, cuBLAS, Vulkan
  • GGUF format parser
  • Llama 3 tokenizer based on minbpe
  • Llama 3 inference with Grouped-Query Attention
  • Support Llama 3.1 (ad-hoc RoPE scaling) and 3.2 (tie word embeddings)
  • Support F16, BF16 weights + Q8_0 and Q4_0 quantizations
  • Simple CLI with --chat and --instruct modes.
  • GraalVM's Native Image support (EA builds here)
  • Java Message bus a la ROSJava

Interactive --chat mode in action:

Setup

Download pure Q4_0 and (optionally) Q8_0 quantized .gguf files from:

The pure Q4_0 quantized models are recommended, except for the very small models (1B), please be gentle with huggingface.co servers:

# Llama 3.2 (3B)
curl -L -O https://huggingface.co/mukel/Llama-3.2-3B-Instruct-GGUF/resolve/main/Llama-3.2-3B-Instruct-Q4_0.gguf

# Llama 3.2 (1B)
curl -L -O https://huggingface.co/mukel/Llama-3.2-1B-Instruct-GGUF/resolve/main/Llama-3.2-1B-Instruct-Q8_0.gguf

# Llama 3.1 (8B)
curl -L -O https://huggingface.co/mukel/Meta-Llama-3.1-8B-Instruct-GGUF/resolve/main/Meta-Llama-3.1-8B-Instruct-Q4_0.gguf

# Llama 3 (8B)
curl -L -O https://huggingface.co/mukel/Meta-Llama-3-8B-Instruct-GGUF/resolve/main/Meta-Llama-3-8B-Instruct-Q4_0.gguf

# Optionally download the Q8_0 quantized models
# curl -L -O https://huggingface.co/mukel/Meta-Llama-3-8B-Instruct-GGUF/resolve/main/Meta-Llama-3-8B-Instruct-Q8_0.gguf
# curl -L -O https://huggingface.co/mukel/Meta-Llama-3.1-8B-Instruct-GGUF/resolve/main/Meta-Llama-3.1-8B-Instruct-Q8_0.gguf

Optional: quantize to pure Q4_0 manually

In the wild, Q8_0 quantizations are fine, but Q4_0 quantizations are rarely pure e.g. the token_embd.weights/output.weights tensor are quantized with Q6_K, instead of Q4_0.
A pure Q4_0 quantization can be generated from a high precision (F32, F16, BFLOAT16) .gguf source with the llama-quantize utility from llama.cpp as follows:

./llama-quantize --pure ./Meta-Llama-3-8B-Instruct-F32.gguf ./Meta-Llama-3-8B-Instruct-Q4_0.gguf Q4_0

Build and run

Java 21+ is required, in particular the MemorySegment mmap-ing feature.

Run from source

java --enable-preview --source 21 --add-modules jdk.incubator.vector org.ros.ROSCore <ROS node class> -i --model Meta-Llama-3-8B-Instruct-Q4_0.gguf

Optional: Makefile + manually build and run

llama.cpp

Vanilla llama.cpp built with make.

./llama-cli --version                                                                                                                                                                          130 ↵
version: 3862 (3f1ae2e3)
built with cc (GCC) 14.2.1 20240805 for x86_64-pc-linux-gnu

Executed as follows:

./llama-bench -m Llama-3.2-1B-Instruct-Q4_0.gguf -p 0 -n 128

Llama3.java

taskset -c 0-15 ./llama3 \
  --model ./Llama-3-1B-Instruct-Q4_0.gguf \
  --max-tokens 128 \
  --seed 42 \
  --stream false \
  --prompt "Why is the sky blue?"

Hardware specs: 2019 AMD Ryzen 3950X 16C/32T 64GB (3800) Linux 6.6.47.

**Notes
Running on a single CCD e.g. taskset -c 0-15 ./llama3 ... since inference is constrained by memory bandwidth.

Results

License

MIT

About

JDK 25 FFI to Llama.cpp to RosJavaLite model runner for AI embodied robotics.

Resources

License

Stars

1 star

Watchers

0 watching

Forks

Releases

No releases published

Packages

 
 
 

Contributors

Languages