tiyuvta.ai’s cover photo
tiyuvta.ai

tiyuvta.ai

Software Development

Open models tuned to your data, running on hardware you control. Cut your LLM cost of goods.

About us

Your API bill is your cost of goods. It grows with every customer you win, and someone else sets the price. tiyuvta moves that workload onto hardware you control — your own cards, a rented GPU server, or dedicated machines we manage — running open models fine-tuned on your schemas, your terminology and your language. Same work, a fraction of the bill. Teams that do this seriously land near a third of what they were paying, on better models than they started with, and the hardware keeps working after it is paid for. WHAT WE DO • Cut the cost of goods on an LLM product: tuned models, tuned serving stack, hardware you own or rent • Agent systems that do real work: customers answered, crisis triage, documents and paperwork, company knowledge that updates daily • Fine-tuning on your data, delivered as weights you keep • BYOC — your cards, our tuning, the numbers proven on your machine • Prepaid OpenAI-compatible inference API, key in minutes. If your workload is small we will say so and sell you nothing else. WHY TEAMS TRUST THIS • Every number carries its build, date and measurement protocol. Lab figures are marked as lab figures. • Public research and reproducible benches: tiyuvta.ai/blog, plus work published on Hugging Face and GitHub • What we deliver, you hold: fine-tuned weights, serving config and runbook, on your hardware. Escrow available on request. • OpenAI-compatible surface and open formats, so moving is a base-URL change. No lock-in. • You talk to Avi Fenesh, the engineer who tunes the stack. No sales queue, no discovery fee, no paid scoping call. START HERE → Book 20 minutes with an engineer: https://calendly.com/avifenesh-tiyuvta/30min → Get an API key in minutes: https://inference.tiyuvta.ai → Read the lab: https://tiyuvta.ai/blog → Or write directly: hello@tiyuvta.ai

Website
https://tiyuvta.ai
Industry
Software Development
Company size
1 employee
Headquarters
Tel Aviv
Type
Self-Employed
Founded
2026
Specialties
LLM inference, LLM fine-tuning, Open-weight models, Self-hosted LLM, GPU serving, vLLM, Agent systems, AI agents, RAG, Model quantization, Speculative decoding, Inference API, LLM cost optimization, On-premise AI, Private AI deployment, Dedicated GPU hosting, MLOps, Generative AI consulting, AI infrastructure, and Hebrew language models

Locations

Employees at tiyuvta.ai

Updates

  • What you might miss if you look at the i/o price only: DeepSeek V4.1 Flash is a smart and efficient model. Compared to the latest leading flash models, its input and output price per 1m tokens is costlier. It is not expensive, 0.3 for 1 m input and 1.2 for 1M output, but if you compare it to GLM 5.3 Flash, very close to DSV4.1F in performance, it costs 0.15/0.5. Seems like less than a half. But the specialty of the model is that it is built to be extremely efficient to cache - and the cache price is only 0.03 per 1m tokens. If you run long sessions in loops, most of what you pay for is cache. If you use a good provider with good inference engineers, cache might be > 90% of what you pay for. Tiyuvta stats over the last week: 93% of the input is cached. Do the math. https://lnkd.in/dCe3bduJ #kvcache #deepseek #paretofrontier

    View organization page for Arena

    29,145 followers

    DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net improvement at $0.07 cost per median task, it’s reshaped the Pareto frontier for Agent Arena! Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost. For comparison, it retains: - 98% of Hy4 preview’s net improvement, at 73% lower cost - 76% of Kimi K3 (Max)’s performance, at 92% lower cost. Against models as powerful as Fable 5 or stronger, DeepSeek-V4.1-Flash (Max) retains 35–54% of their net improvement at 97–99% lower cost. Those top models cost 37–76× more per task. Net improvement over Arena baseline | Median cost/task - Claude Fable 5.1 (Max): +13.90% | $4.54 - GPT 6 Astra (Max): +11.90% | $4.09 - Claude Opus 5 (Max): +11.09% | $3.52 - Claude Opus 5 (High): +10.49% | $2.24 - Claude Fable 5 (High): +9.03% | $2.19 - Claude Opus 4.8 (High): +7.75% | $1.36 - GPT 5.6 Sol (xHigh): +7.40% | $1.09 - Kimi K3 (Max): +6.39% | $0.77 - Hy4 preview: +4.96% | $0.22 - DeepSeek-V4.1-Flash (Max): +4.87% | $0.06 With this release, GPT-5.6 Luna (xHigh), GLM-5.3-Flash, and DeepSeek-V4-Flash fell off the Pareto frontier for Agent Arena. Congrats again to the DeepSeek AI team on this release!

  • tiyuvta.ai reposted this

    View organization page for Arena

    29,145 followers

    DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net improvement at $0.07 cost per median task, it’s reshaped the Pareto frontier for Agent Arena! Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost. For comparison, it retains: - 98% of Hy4 preview’s net improvement, at 73% lower cost - 76% of Kimi K3 (Max)’s performance, at 92% lower cost. Against models as powerful as Fable 5 or stronger, DeepSeek-V4.1-Flash (Max) retains 35–54% of their net improvement at 97–99% lower cost. Those top models cost 37–76× more per task. Net improvement over Arena baseline | Median cost/task - Claude Fable 5.1 (Max): +13.90% | $4.54 - GPT 6 Astra (Max): +11.90% | $4.09 - Claude Opus 5 (Max): +11.09% | $3.52 - Claude Opus 5 (High): +10.49% | $2.24 - Claude Fable 5 (High): +9.03% | $2.19 - Claude Opus 4.8 (High): +7.75% | $1.36 - GPT 5.6 Sol (xHigh): +7.40% | $1.09 - Kimi K3 (Max): +6.39% | $0.77 - Hy4 preview: +4.96% | $0.22 - DeepSeek-V4.1-Flash (Max): +4.87% | $0.06 With this release, GPT-5.6 Luna (xHigh), GLM-5.3-Flash, and DeepSeek-V4-Flash fell off the Pareto frontier for Agent Arena. Congrats again to the DeepSeek AI team on this release!

  • If your product runs on an LLM, the API bill is your cost of goods. It grows with every customer you win, and you do not control the price. We move that work onto hardware you control at self-hosting cost or on secure servers in the cloud, which you can have with the option for us to manage the servers for you on our hardware: models tuned to your data and your language, the surrounding pipeline, and a serving stack tuned to your workload. Cutting a token bill to a fraction of itself is the whole point. Teams doing this seriously land near a third of what they were paying on better models than they started with, and the hardware keeps working after it is paid for. The same work also runs the internal side: customers are answered, documents are handled, and company knowledge is kept current. And if your workload is small, start on our API in minutes instead. We are the same people either way. https://tiyuvta.ai/

    • self host, api, llm, ai
  • DeepSeek 4.1 Flash is now supported on Tiyuvta. What makes this model interesting to us is the inference story. DeepSeek is pushing toward a model that delivers very high capability without requiring the entire model to be active for every token. That changes the economics of serving: less compute, less memory pressure, and a much more practical path to running a very large model efficiently. This is exactly the direction we care about at Tiyuvta. Not just better models - better models that can actually be served efficiently. You can now try DeepSeek 4.1 Flash on Tiyuvta inference via API. And if you try it, like it, and your business actually needs this class of model at a much lower serving cost, we can help you run it on infrastructure dedicated to your workload as well. Leave a comment or email us. hello@tiyuvta.ai https://lnkd.in/dVGZuPDQ

    • No alternative text description for this image
  • If you don't use agents for coding, the SWE benches on the model's front page are not what you should care about. Most SOTA models are racing in coding and math performance. If your agent is supposed to handle customer reports, there are plenty of good models that cost 1/10 that will do the same job with the same results. Check the right numbers.

    • No alternative text description for this image
  • זה קל להבין למה להחזיק AI שרץ In house אם המוצר שלכם מבוסס AI. זאת העלות העיקרית. אבל נניח ואתם לא "AI product". אתם רק משתמשים לכמה דברים. אם אתם נותנים לכל עובד seat בקורסור, אז לא על זה אני מדבר. ובלי קשר זה לא הפתרון שבטוח תפור עליכם וזה מה שאתם צריכים, אבל לחלקחים כן. יש לכם בוט של שירות לקוחות, הוא עובד 18 שעות ביום וחוסך לכם מלא שעות אדם בפניות שנפתרות או מופנות למוקד הנכון. יש לכם אייג'נט שמרכז את המידע של החברה, הוא בונה מערכת זיכרון, הוא מנקה את הקלאשינג בין העובדות הסותרות, הוא מכין את הקרקע לבוטים, שיענו נכון. אחרון, אייג'נט שמכיר את החברה ומחפש לכם לידים, הרבה tools, mcp, והרבה גישה למידע שלכם, in and out. אתם משתמשים במודל זול אבל של חברה מובילה, כי לא תשימו את העסק שלכם אצל איזשהי מעבדת no-name. מחיר חודשי, בלי שימוש כבד מדי - 10k שח בחודש. סביר ושווה את זה. היום מערכת כוללת של 5090 עולה לכם בסביבות ה40k שח. עם memra אתם מריצים את qwen 3.8 24b בקצב של 111 טוקן בשנייה, עם זמן תגובה של פחות מחצי שנייה. אתם יכולים להחזיק במקביל 3 סשנים כדי למקבל את כל האייג'נטים שלכם. רק כדי להבין, כשאתם משלמים על "fast" אתם מקבלים אולי 70 טוקן לשנייה, יש לכם את הזמן שזה רץ ברשת ולוקח הרבה יותר מחצי שנייה לטוקן ראשון. על ה40k האלו את גם מקבלים כוללים במס פחת של 33% במשך שלוש שנים אז המחיר הוא תכלס 30k. נניח שזה באמת רק לשלוש שנים, אתם ב1100 לחודש. במקום 10k. תוסיפו את העלות שלנו, שתלויה בהמון דברים ותלויה במה שאתם צריכים, אבל על גבי 3 שנים זה יגיע מקסימום לחצי. אם זה יותר מחצי אנחנו כנראה נציע לכם פתרון אחר. בונוס: אם העסק שלכם מתפוצץ, אתם פשוט מוסיפים עוד כרטיס, 20k, לאותה מערכת. לא מכפילים את המחיר. ואם הקמתם איתנו, אנחנו מייעצים ומלווים, ומסכמים מראש כמה אחוז מהוצאת ההקמה יעלה להגדיל, להחליף מודל, וכו. בונוס 2: אייג'נטים זה ההתמחות שלנו מהבית, על הדרך אנחנו ניתן טיפים על איך לחסוך, איך לבנות משהו שלא עושה פאדיחות, איזה מודל באמת חזק מספיק, ונצייר לכם דיאגרמה של הפייפליין שלכם סתם כי זה כיף לנו. אם בא לכם לשמוע עוד תשאירו תגובה, שלחו הודעה או תשאירו מייל ב hello@tiyuvta.ai

  • עלתה גרסה בעברית לאתר. שירותים עם דוגמאות אמיתיות: fine-tune, BYOC, דיפלוימנט על השרתים שלנו עם המודלים שנעזור לכם להתאים לסרוויס שלכם, וapi זול ופשוט מהיר יותר. מה שמתאים *לכם*. רוצים לדבר? תשאירו פרטים בתגובה או מייל ל hello@tiyuvta.ai. https://lnkd.in/dSWt2XtF #בינהמלאכותית #הייטק #סטארטאפ

    • No alternative text description for this image
  • Most company writing is not a blank page. It is an email that came out too sharp. Notes nobody shaped. A brief that another system needs as JSON. We published a use case for that: write the work, on DeepSeek-V4.1-Flash. Rewrite the email. Turn notes into decisions and actions. Turn a rough brief into structured JSON your app can save. https://lnkd.in/dCD4qBzi

    • No alternative text description for this image

Similar pages