aiha-lab
Popular repositories Loading
-
InfiniPot-V
InfiniPot-V Public[NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
-
Attention-Head-Pruning
Attention-Head-Pruning PublicLayer-wise Pruning of Transformer Heads for Efficient Language Modeling
Repositories
- tangram Public
A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM
- tangram-page Public
An Unstructured and Memory-Efficient Framework for LLM Serving and KV Cache Management.
- BeaconKV Public
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
- OSWorld-wkl Public Forked from xlang-ai/OSWorld
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- InfiniPot-V Public
[NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
- MapCoder-Lite Public
[EACL2026 Findings] MapCoder-Lite: Distilling Multi-Agent Coding into a Single Small LLM
- NPU-Based-Real-Time-Blind-Streaming-Assistant Public
A real-time streaming assistant powered by Rebellions NPU, designed to operate without visual feedback and optimized for low-latency.
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…