HPC Documentation
Minerva Generative Assistant Tool: Supported AI Models
Model Catalog
This page provides a reference for all Large Language Models (LLMs) deployed on the Minerva HPC Generative AI Assistant infrastructure. All models run on local GPU clusters (NVIDIA H100s) ensuring institutional privacy with zero external telemetry.
| Model | Parameters | Best For | More Information |
|---|---|---|---|
gemma3:4b |
4B | Lightweight multimodal model for fast text and image tasks | Google Gemma 3 |
gemma3:12b-it-fp16 |
12B | Full-precision Gemma 3 model for high-quality text and vision tasks | Google Gemma 3 |
gemma3:12b-it-q8_0 |
12B | High-precision 8-bit quantized Gemma 3 model for multimodal reasoning | Google Gemma 3 |
gemma3:12b-it-q4_K_M |
12B | Balanced 4-bit quantized Gemma 3 12B for resource-efficient performance | Google Gemma 3 |
gemma3:12b-it-qat |
12B | Quantization-Aware Trained 12B model balancing low memory and high quality | Google Gemma 3 |
gemma3:27b-it-fp16 |
27B | Full-precision high-capacity model for complex multimodal reasoning | Google Gemma 3 |
gemma3:27b-it-q8_0 |
27B | 8-bit quantized Gemma 3 27B for precise instruction following & vision tasks | Google Gemma 3 |
gemma3:27b-it-q4_K_M |
27B | 4-bit quantized Gemma 3 27B offering strong reasoning with low VRAM footprint | Google Gemma 3 |
gemma3:27b-it-qat |
27B | Quantization-Aware Trained 27B model delivering near-BF16 precision | Google Gemma 3 |
gemma4:26b |
26B | High-performance Gemma 4 model for deep reasoning and analytical tasks | Google Gemma 4 |
gemma4:31b |
31B | Next-gen Gemma 4 model for advanced multimodal and complex logic workflows | Google Gemma 4 |
gpt-oss:20b |
21B | Open-weight model optimized for fast reasoning and function calling | OpenAI GPT-OSS |
gpt-oss:120b |
117B | Large open-weight model for enterprise-grade reasoning and agent workflows | OpenAI GPT-OSS |
llama3.3:70b |
71B | Best-in-class reasoning, handles complex queries, lowest hallucination rate | Meta Llama 3.3 |
llama4:latest |
109B | Flagship Mixture-of-Experts multimodal model for complex reasoning and chat | Meta Llama 4 |
llava:7b |
7B | Multimodal vision-language model (image + text) | LLaVA GitHub |
medgemma1.5:latest |
4B | Domain-specific clinical reasoning and medical text understanding | Google MedGemma |
muse-glimmer:latest |
28B | Agentic model optimized for multi-step reasoning, tool use, and long-horizon tasks | Meta Muse Glimmer |
nemotron-3-nano:30b |
32B | GPU-optimized architecture for fast enterprise language generation | NVIDIA Nemotron |
phi4:latest |
15B | High-efficiency small model with strong reasoning, math, and coding skills | Microsoft Phi-4 |
tinyllama:latest |
1B | Lightweight, fast inference with minimal resource footprint | TinyLlama GitHub |
Note: LLM models available in the Assistant tool may make mistakes. Always review outputs carefully before using them.
