Model Catalog

This page provides a reference for all Large Language Models (LLMs) deployed on the Minerva HPC Generative AI Assistant infrastructure. All models run on local GPU clusters (NVIDIA H100s) ensuring institutional privacy with zero external telemetry.

Model Parameters Best For More Information
gemma3:4b 4B Lightweight multimodal model for fast text and image tasks Google Gemma 3
gemma3:12b-it-fp16 12B Full-precision Gemma 3 model for high-quality text and vision tasks Google Gemma 3
gemma3:12b-it-q8_0 12B High-precision 8-bit quantized Gemma 3 model for multimodal reasoning Google Gemma 3
gemma3:12b-it-q4_K_M 12B Balanced 4-bit quantized Gemma 3 12B for resource-efficient performance Google Gemma 3
gemma3:12b-it-qat 12B Quantization-Aware Trained 12B model balancing low memory and high quality Google Gemma 3
gemma3:27b-it-fp16 27B Full-precision high-capacity model for complex multimodal reasoning Google Gemma 3
gemma3:27b-it-q8_0 27B 8-bit quantized Gemma 3 27B for precise instruction following & vision tasks Google Gemma 3
gemma3:27b-it-q4_K_M 27B 4-bit quantized Gemma 3 27B offering strong reasoning with low VRAM footprint Google Gemma 3
gemma3:27b-it-qat 27B Quantization-Aware Trained 27B model delivering near-BF16 precision Google Gemma 3
gemma4:26b 26B High-performance Gemma 4 model for deep reasoning and analytical tasks Google Gemma 4
gemma4:31b 31B Next-gen Gemma 4 model for advanced multimodal and complex logic workflows Google Gemma 4
gpt-oss:20b 21B Open-weight model optimized for fast reasoning and function calling OpenAI GPT-OSS
gpt-oss:120b 117B Large open-weight model for enterprise-grade reasoning and agent workflows OpenAI GPT-OSS
llama3.3:70b 71B Best-in-class reasoning, handles complex queries, lowest hallucination rate Meta Llama 3.3
llama4:latest 109B Flagship Mixture-of-Experts multimodal model for complex reasoning and chat Meta Llama 4
llava:7b 7B Multimodal vision-language model (image + text) LLaVA GitHub
medgemma1.5:latest 4B Domain-specific clinical reasoning and medical text understanding Google MedGemma
muse-glimmer:latest 28B Agentic model optimized for multi-step reasoning, tool use, and long-horizon tasks Meta Muse Glimmer
nemotron-3-nano:30b 32B GPU-optimized architecture for fast enterprise language generation NVIDIA Nemotron
phi4:latest 15B High-efficiency small model with strong reasoning, math, and coding skills Microsoft Phi-4
tinyllama:latest 1B Lightweight, fast inference with minimal resource footprint TinyLlama GitHub
Note: LLM models available in the Assistant tool may make mistakes. Always review outputs carefully before using them.