GitHub - krishkc5/energy_aware_quantization: End-to-end framework for evaluating accuracy, latency, and energy of quantized LLMs (FP32, FP16, INT8). Includes pre-tokenized SST-2 datasets, reproducible inference harness, GPU power logging via nvidia-smi, and analysis tools for energy-aware model deployment.

Branches Tags

Name		Name	Last commit message	Last commit date
Latest commit History 98 Commits
datasets		datasets
notebooks		notebooks
.gitignore		.gitignore
eaq_presentation.pdf		eaq_presentation.pdf
eaq_report.pdf		eaq_report.pdf
requirements.txt		requirements.txt

About

End-to-end framework for evaluating accuracy, latency, and energy of quantized LLMs (FP32, FP16, INT8). Includes pre-tokenized SST-2 datasets, reproducible inference harness, GPU power logging via nvidia-smi, and analysis tools for energy-aware model deployment.