You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models.
a drop-in replacement fork of the original bincode crate: "A binary serialization / deserialization strategy for transforming structs into bytes and vice versa!"
🛠️ Implement TOON in PHP for efficient serialization of JSON-like data, optimizing parsing for Large Language Models while maintaining clarity and structure.
Lsglang is a special extension of sglang that fully utilizes CPU and GPU computing resources with an efficient GPU parallel + NUMA parallel architecture, suitable for MOE model hybrid inference.
Schemaless Protobuf inspection toolkit for .NET. Decodes raw wire-format payloads with no .proto files, generated classes, or schemas - auto-intercepting HTTP and gRPC (unary + streaming) traffic into human-readable trees or JSON for diagnostics, auditing, and reverse engineering.