xJoePec commited on
Commit
ac75d74
·
verified ·
1 Parent(s): 7cccd0c

Upload MODEL_CARD.md

Browse files
Files changed (1) hide show
  1. MODEL_CARD.md +257 -0
MODEL_CARD.md ADDED
@@ -0,0 +1,257 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Model Card: Galena-2B (Granite 3.3 Math & Physics)
2
+
3
+ ## Model Description
4
+
5
+ **Galena-2B** is a specialized 2-billion parameter language model optimized for mathematical reasoning and physics problem-solving. It is derived from IBM's Granite 3.3-2B Instruct base model through parameter-efficient fine-tuning (LoRA) on curated datasets focused on advanced calculations and physics concepts.
6
+
7
+ - **Developed by:** [Your Name/Organization]
8
+ - **Base Model:** [IBM Granite 3.3-2B Instruct](https://huggingface.co/ibm-granite/granite-3.3-2b-instruct)
9
+ - **Model Type:** Causal Language Model (Decoder-only Transformer)
10
+ - **Language:** English
11
+ - **License:** Apache 2.0
12
+ - **Fine-tuned from:** ibm-granite/granite-3.3-2b-instruct
13
+
14
+ ## Model Architecture
15
+
16
+ - **Architecture:** GraniteForCausalLM
17
+ - **Parameters:** 2.0B
18
+ - **Layers:** 40
19
+ - **Hidden Size:** 2048
20
+ - **Attention Heads:** 32 (query) / 8 (key-value, GQA)
21
+ - **Intermediate Size:** 8192
22
+ - **Vocabulary Size:** 49,159 tokens
23
+ - **Context Window:** 131,072 tokens (128k)
24
+ - **Precision:** bfloat16 (training & inference)
25
+ - **Activation Function:** SiLU (Swish)
26
+
27
+ ### Key Features
28
+
29
+ - **Grouped Query Attention (GQA)** for efficient inference
30
+ - **RoPE Embeddings** with extended context support (theta=10M)
31
+ - **Attention & Logits Scaling** for training stability
32
+ - **Embedding Multiplier** (12.0) and Residual Multiplier (0.22)
33
+
34
+ ## Intended Use
35
+
36
+ ### Primary Use Cases
37
+
38
+ - **Educational Applications:** Teaching and learning advanced mathematics and physics
39
+ - **Research Tools:** Assisting with physics problem formulation and mathematical reasoning
40
+ - **Conversational AI:** Domain-specific chatbots for STEM topics
41
+ - **Tool-Augmented Reasoning:** Integration with calculators and symbolic math engines
42
+
43
+ ### Out-of-Scope Use
44
+
45
+ - **Critical Decision Making:** Not suitable for medical, legal, or safety-critical applications
46
+ - **General-Purpose Conversational AI:** Optimized for math/physics; may underperform on general topics
47
+ - **Production Systems:** This is a research/educational model without production guarantees
48
+ - **Factual Information Retrieval:** May hallucinate; always verify outputs
49
+
50
+ ## Training Data
51
+
52
+ The model was fine-tuned on a carefully curated dataset of 26,000 instruction-response pairs blending two specialized datasets:
53
+
54
+ ### 1. NVIDIA Nemotron-RL-Math (Advanced Calculations)
55
+
56
+ - **Source:** `nvidia/Nemotron-RL-math-advanced_calculations`
57
+ - **Content:** Complex mathematical problems with step-by-step reasoning traces
58
+ - **Features:** Tool-augmented reasoning, calculator integration, multi-step problem decomposition
59
+ - **Format:** Instruction-following with detailed solution traces
60
+
61
+ ### 2. CAMEL-AI Physics Dataset
62
+
63
+ - **Source:** `camel-ai/physics`
64
+ - **Content:** Physics dialogue pairs covering diverse topics and subtopics
65
+ - **Features:** Conceptual explanations, problem-solving, physics principles
66
+ - **Metadata:** Topic and subtopic categorization for structured learning
67
+
68
+ ### Data Preparation
69
+
70
+ - **Preprocessing:** `scripts/prepare_math_physics.py` in parent GRANITE repository
71
+ - **Format Conversion:** Unified into Granite's chat format (`<|user|>`/`<|assistant|>` tags)
72
+ - **Output:** `data/math_physics.jsonl` (26k examples)
73
+ - **Token Length:** Max sequence length capped at 512 tokens during training
74
+
75
+ ## Training Procedure
76
+
77
+ ### Training Hyperparameters
78
+
79
+ - **Method:** QLoRA (Quantized Low-Rank Adaptation)
80
+ - **Base Model Precision:** 4-bit quantization (NF4)
81
+ - **LoRA Rank:** Default (typically 8-16)
82
+ - **LoRA Alpha:** Default
83
+ - **Target Modules:** Query, Key, Value, Output projections
84
+ - **Gradient Checkpointing:** Enabled
85
+ - **Mixed Precision:** bfloat16
86
+
87
+ ### Training Configuration
88
+
89
+ ```python
90
+ {
91
+ "base_model": "ibm-granite/granite-3.3-2b-instruct",
92
+ "dataset_path": "data/math_physics.jsonl",
93
+ "output_dir": "outputs/granite-math-physics-lora",
94
+ "use_4bit": true,
95
+ "per_device_train_batch_size": 1,
96
+ "gradient_accumulation_steps": 4,
97
+ "effective_batch_size": 4,
98
+ "num_train_epochs": 1,
99
+ "max_steps": 500,
100
+ "max_seq_length": 512,
101
+ "learning_rate": "2e-4 (default)",
102
+ "batching_strategy": "padding",
103
+ "optimizer": "paged_adamw_8bit",
104
+ "bf16": true
105
+ }
106
+ ```
107
+
108
+ ### Training Infrastructure
109
+
110
+ - **Hardware:** NVIDIA GeForce RTX 4060 (8GB VRAM)
111
+ - **Software Stack:**
112
+ - PyTorch 2.x
113
+ - Hugging Face Transformers 4.44+
114
+ - PEFT 0.11+
115
+ - bitsandbytes 0.43+
116
+ - CUDA 12.1
117
+ - **Training Time:** ~500 steps (1 epoch over 26k examples with batch size 4)
118
+ - **Checkpointing:** LoRA adapters saved every N steps
119
+
120
+ ### Post-Training
121
+
122
+ 1. **Adapter Merging:** LoRA adapters merged back into base weights using `scripts/merge_lora.py`
123
+ 2. **GGUF Conversion:** Exported to F16 GGUF format via `llama.cpp/convert_hf_to_gguf.py`
124
+ 3. **Formats Produced:**
125
+ - Hugging Face Transformers (safetensors)
126
+ - GGUF F16 (llama.cpp compatible)
127
+
128
+ ## Evaluation
129
+
130
+ ### Qualitative Assessment
131
+
132
+ The model demonstrates improved performance on:
133
+
134
+ - Multi-step mathematical reasoning
135
+ - Physics problem explanation
136
+ - Calculator-augmented computation tasks
137
+ - Domain-specific terminology and notation
138
+
139
+ ### Limitations
140
+
141
+ - **Limited Training Steps:** Only 500 training steps; longer training may improve performance
142
+ - **Domain Specialization:** May sacrifice general capabilities for math/physics expertise
143
+ - **Hallucination Risk:** Can generate plausible but incorrect solutions
144
+ - **Tool Integration:** Expects calculator tools in reasoning traces; standalone performance may vary
145
+ - **Context Window:** Fine-tuned on 512-token sequences; full 128k context not extensively tested
146
+
147
+ ## Bias, Risks, and Limitations
148
+
149
+ ### Known Limitations
150
+
151
+ 1. **Domain Specificity:** Optimized for math/physics; general knowledge may be limited
152
+ 2. **Factual Accuracy:** No guarantee of correctness; outputs should be verified
153
+ 3. **Training Data Bias:** Inherits biases from Nemotron and CAMEL-AI datasets
154
+ 4. **Base Model Limitations:** Retains all limitations of Granite 3.3-2B Instruct
155
+ 5. **Small Training Set:** 26k examples may not cover all edge cases
156
+
157
+ ### Ethical Considerations
158
+
159
+ - **Educational Use:** Should supplement, not replace, human instruction
160
+ - **Verification Required:** Always validate mathematical and scientific outputs
161
+ - **Accessibility:** May use technical jargon inaccessible to beginners
162
+ - **Dataset Provenance:** Users should review source dataset licenses and terms
163
+
164
+ ### Recommendations
165
+
166
+ - Use as an educational aid, not a source of truth
167
+ - Implement output validation for critical applications
168
+ - Combine with symbolic computation tools for verification
169
+ - Monitor for hallucinations and incorrect reasoning
170
+ - Consider fine-tuning on domain-specific data for production use
171
+
172
+ ## Environmental Impact
173
+
174
+ - **Hardware:** NVIDIA RTX 4060 (8GB VRAM)
175
+ - **Training Duration:** ~500 steps (estimated 1-2 hours)
176
+ - **Energy Consumption:** Estimated <1 kWh for training
177
+ - **Carbon Footprint:** Minimal due to efficient LoRA training
178
+
179
+ ## Technical Specifications
180
+
181
+ ### Model Formats
182
+
183
+ | Format | Precision | Size | Compatible Frameworks |
184
+ |--------|-----------|------|-----------------------|
185
+ | Hugging Face Transformers | bfloat16 | ~5.0 GB | PyTorch, Transformers, vLLM, TGI |
186
+ | GGUF F16 | float16 | ~4.7 GB | llama.cpp, Ollama, LM Studio |
187
+
188
+ ### System Requirements
189
+
190
+ **Minimum (CPU Inference):**
191
+ - RAM: 8 GB
192
+ - Storage: 10 GB free space
193
+ - CPU: Modern x86-64 with AVX2 support
194
+
195
+ **Recommended (GPU Inference):**
196
+ - GPU: 6+ GB VRAM (RTX 3060, A4000, or better)
197
+ - RAM: 16 GB
198
+ - CUDA 12.1+ or ROCm 5.7+
199
+
200
+ ### Loading & Inference
201
+
202
+ Before running inference, pull the artifacts into `models/math-physics/`:
203
+
204
+ ```bash
205
+ python scripts/download_artifacts.py --artifact all
206
+ ```
207
+
208
+ **Transformers (Python):**
209
+ ```python
210
+ from transformers import AutoModelForCausalLM, AutoTokenizer
211
+
212
+ model = AutoModelForCausalLM.from_pretrained(
213
+ "models/math-physics/hf",
214
+ device_map="auto",
215
+ trust_remote_code=True
216
+ )
217
+ tokenizer = AutoTokenizer.from_pretrained("models/math-physics/hf")
218
+ ```
219
+
220
+ **llama.cpp (Command Line):**
221
+ ```bash
222
+ ./llama-cli -m granite-math-physics-f16.gguf -p "Your prompt" -n 256
223
+ ```
224
+
225
+ ## Citation
226
+
227
+ ```bibtex
228
+ @software{galena_2b_2024,
229
+ title = {Galena-2B: Granite 3.3 Math & Physics Model},
230
+ author = {Your Name},
231
+ year = {2024},
232
+ url = {https://github.com/yourusername/galena-2B},
233
+ note = {Fine-tuned from IBM Granite 3.3-2B Instruct on math and physics datasets}
234
+ }
235
+ ```
236
+
237
+ ## Acknowledgments
238
+
239
+ - IBM Research for the Granite 3.3 foundation model
240
+ - NVIDIA for the Nemotron-RL-Math dataset
241
+ - CAMEL-AI for the physics dialogue dataset
242
+ - Hugging Face for training infrastructure and libraries
243
+
244
+ ## Contact
245
+
246
+ For questions, issues, or contributions:
247
+ - **Repository:** [GitHub Issues](https://github.com/yourusername/galena-2B/issues)
248
+ - **Email:** your.email@example.com
249
+
250
+ ## Changelog
251
+
252
+ ### Version 1.0 (2024-11-17)
253
+
254
+ - Initial release
255
+ - Fine-tuned on 26k math/physics examples
256
+ - 500 training steps with QLoRA
257
+ - Hugging Face and GGUF formats released