Skip to content
View avifenesh's full-sized avatar
🦥
Just hanging around
🦥
Just hanging around

Highlights

  • Pro

Organizations

@aws @amazon-contributing @valkey-io

Block or report avifenesh

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
avifenesh/README.md

Hey, I'm Avi 👋

I like machines close to the metal and results you can measure. I run tiyuvta, a research lab that helps companies choose, adapt and deploy open models on their hardware or cloud account. I also build in-memory data systems at AWS ElastiCache, and maintain open-source tools people actually run.

Research notes and working papers live at avifenesh.ai. Issues, questions, and counterexamples are always welcome.

tiyuvta

tiyuvta is my research lab. We help companies run open models on their hardware or cloud account. The work includes model and hardware assessment, deployment and serving-engine optimization, fine-tuning and evaluation, and cloud setup. We compare options on your workload and budget, then agree the deliverables, handover and support.

Our research in model execution, quantization, pruning, speculative decoding and Hebrew models informs that work. See the services and discuss your project.

The lab's models are on Hugging Face: NVFP4 builds of GLM, DeepSeek, MiMo, Qwen and DictaLM, MTP draft heads in GGUF, and Phonon-2 in ONNX.

ML research

  • Running now at tiyuvta: a negative-eigenvalue gate retrofit for Gated DeltaNet hybrids, +48 points at 4B and +43 at 9B on state tracking at 4x the training length, now scaling to Qwen3.8-27B (record). And Layby: KV cache placement for idle LLM sessions by when they come back. A cost rule with no fitted constants and Layby-Dwell, an open 421M return-time model; adapters for vLLM and SGLang and the open ReturnBench benchmark. On a slow disk with Qwen3-8B it cuts p95 time to first token to 0.59 of an LRU CPU tier with the median unchanged, half of LMCache's best setting; with GLM 5.3 Flash on a shared disk, 0.79 at p95 with p99 held at the baseline, where writing everything to disk makes p95 1.35 times worse. Paper on arXiv to follow. Details at avifenesh.ai/research.
  • memra: tiyuvta's from-scratch Rust + CUDA research engine for RTX PRO 6000 Blackwell and RTX 5090, an instrument behind the studies below. Safetensors is the tuned path, GGUF stays supported, and a mechanism that wins on one card and loses on the other becomes a per-device default rather than a compromise. On crates.io with prebuilt binaries.
  • hqmtp: MTP draft-head lab, concluded. Function cuts (pruning, low-rank, distillation) pay a 10–19-point off-distribution tax that fidelity cuts don't; the zero-training trimmed-vocabulary recipe won at 1.8–2.7× end to end. The negative results stay in the ledger.
  • recipe-lab: layer-loop weight sharing + ε=λ/(N√L) residual scaling, combined for the first time and tested from zero in 11 pre-registered rounds. In the data-constrained regime the looped model beat FLOPs-matched vanilla in all three mixer families (attention, pure SSM, and hybrid; seven paired runs, zero sign flips), with 26–34% fewer parameters. Rule isolated: loop the state-mixer, never the retriever.
  • mem-retrofit: grafted a product-key memory layer onto a stock dense 4B and ran it against LoRA over sequential updates. The retrofit is free at lr/10; the published forgetting advantage failed 12/12 confidence intervals.
  • More studies with receipts: fixed-compute-frontier (a preregistered kill-gate ledger, ~84 theory lanes) · gemma-expert-atlas (26B MoE expert surgery, 3,840 experts traced) · block-routed-swiglu (near-free kernel, refuted capability) · moe-lab · assumption-excavator.
  • Working papers: small-vocabulary MTP heads · prune, heal, quantize. Methods, failed arms, and evidence in the open.
  • In review upstream:

Maintaining

  • Valkey and its ecosystem. I maintain Valkey GLIDE, the official multi-language client (Rust core, Java/JNI, Node/N-API), and valkey-skills, the official AI skills for the ecosystem, which I started. I contribute to Valkey itself, and a good part of my open-source time goes to the people around it: the clients team, talks, and helping the people who run it.
  • agent-sh, my org: an ecosystem of tools for agent-assisted development, working across Claude Code, Codex, OpenCode, Cursor, and Kiro.
  • glide-mq: Node.js queue on Valkey Streams with a Rust N-API core, plus adapters for Hono, Fastify, Hapi, NestJS and a dashboard.
  • Also around: RustOwl (runtime, memory, and CI work).

Systems

  • Valkey: core contributor; sync-from-replica replication in review. On the side: CRIU copy-on-write live-migration research, under 50 ms of freeze while migrating a 200 GB loaded instance.
  • SGLang router on Valkey: a shared, restart-safe placement index for sgl-kv-indexer, then an event log with worker replay and indexer failover. Measured under load in sglang-valkey-demo.
  • ferrings: io_uring TCP transport for Node.js with a Rust N-API core. 2.5x Node http throughput and 37-57% fewer syscalls per connection in the published bench. On npm.
  • FlowFabric: durable-execution engine in Rust for Valkey, Postgres, and SQLite: lease-safe workers, waitpoints, budgets.
  • layout-audit: DWARF memory-layout analysis: padding, layout diffs, size budgets for C/C++/Rust/Go.
  • scrump: format-aware secret scrubber for binary capture artifacts: perf.data, core dumps, nsys traces, JFR.
  • ocaml-valkey: OCaml 5 + Eio Valkey client, published on opam.

Agents

  • agnix: linter and language server for AI agent configs: 458 rules with autofixes, a GitHub Action, an MCP server, and editor plugins.
  • computer-use-linux / agent-workspace-linux: Linux desktop control over MCP, and isolated agent-owned desktops so an agent never has to touch your real machine.
  • parlar: voice mode for Claude Code and Codex. You talk to a running session, idle or mid-work, and it answers out loud. Local CPU only: VAD, Phonon-2 recognition, Kokoro speech. On npm and crates.io.
  • agentsys: the agent-sh plugin set (workflow, review, ship, deslop, perf) for Claude Code, OpenCode, Codex, Cursor, and Kiro.
  • research-skills: two AgentSkills for ML systems research: write and check the arXiv paper (paper_lint.py), and build, audit and release the code, benchmark and model repo behind it (repo_audit.py).
  • revuto: local PR reviewer that works with any model and learns each repo from its PR history and maintainer feedback. It reviews my own repos.
  • harness tools: read, write, grep, glob, bash, webfetch, lsp and skill tools built for LLM callers, as @agent-sh/harness-* on npm with Rust ports at parity.
  • linubot: Linux desktop app for a team of AI helpers: per-bot providers, memory, and isolated computer workspaces.
  • Codex Desktop for Linux: contributor (120+ commits) to the unofficial Linux build of the ChatGPT/Codex desktop app. Linux Computer Use, Read Aloud and conversation mode, the in-app updater, launcher hardening.

Elsewhere

If something here saved you time, sponsoring helps me keep doing it.

Pinned Loading

  1. agent-sh/computer-use-linux agent-sh/computer-use-linux Public

    Linux desktop control over MCP — AT-SPI, GNOME Shell, Wayland portals, ydotool

    Rust 671 75

  2. agent-sh/agnix agent-sh/agnix Public

    The missing linter and lsp for AI coding assistants. Validate CLAUDE.md, AGENTS.md, SKILL.md, hooks, MCP. Plugin for all major IDEs included, with autofixes.

    Rust 446 34

  3. glide-mq glide-mq Public

    High-performance message queue for Node.js — Valkey/Redis Streams with Rust-native NAPI bindings

    TypeScript 93 2

  4. agent-sh/agentsys agent-sh/agentsys Public

    AI writes code. This automates everything else · 24 plugins · 49 agents · 44 skills · for Claude Code, OpenCode, Codex, Cursor, Kiro.

    JavaScript 995 117

  5. memra memra Public

    Rust + CUDA LLM inference engine for Blackwell (Tuned specifically on RTX PRO 6000, RTX 5090, B200): OpenAI-compatible (+converse and ant) serving, per-model X hardware exactness gates. NVFP4/mixed…

    OpenEdge ABL 2 1

  6. agent-sh/parlar agent-sh/parlar Public

    Voice conversation mode for coding agents: talk to a running Claude Code or Codex session and it talks back

    Rust 2