Projects
-
SpaLLM - Unified compressive adaptation of LLMs with sketching.
Fine-tune compressed models directly using parameter-sharing sketches.
-
SpartanServe - Fast concurrent LLM adapter serving with structurally sparse adapters,
using custom Triton kernels and CUDA graphs.
-
LLM Finetuning Project - Evaluated PEFT methods and 4-bit quantization
for aligning small LLMs (Falcon, Gemma, Phi-2) on domain-specific tasks.
-
NoSQL Document Database - A network accessible document database written in Golang
with RESTful queries, updates, and subscriptions.