DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Your LLM judge gives a different answer on re-runs. How do you test with it?

Your LLM judge gives a different answer on re-runs. How do you test with it?

4
Comments 21
2 min read
Claude Code's ultrareview vs Dromeas Code Review with LLM council

Claude Code's ultrareview vs Dromeas Code Review with LLM council

Comments
2 min read
We ran 15 prompt-injection attacks against a stock local LLM. It failed 3. Here's the free tool we built to check yours.

We ran 15 prompt-injection attacks against a stock local LLM. It failed 3. Here's the free tool we built to check yours.

Comments
3 min read
Don't trust model quotes; use anchors instead

Don't trust model quotes; use anchors instead

Comments
2 min read
GitHub's Canvas Pattern: Why Agentic Workflows Need Spatial State Beyond Chat Scrollback

GitHub's Canvas Pattern: Why Agentic Workflows Need Spatial State Beyond Chat Scrollback

Comments
6 min read
What If AI Is Just Telling You What You Want to Hear?

What If AI Is Just Telling You What You Want to Hear?

Comments
2 min read
Kielo: OpenRouter for Open-Source Models, Powered by Decentralized Compute

Kielo: OpenRouter for Open-Source Models, Powered by Decentralized Compute

Comments
4 min read
Prompt Injection in Copilot Chatbot — Phishing via Client-Controlled Context

Prompt Injection in Copilot Chatbot — Phishing via Client-Controlled Context

Comments
5 min read
Stop Stuffing Your Context Window: 6 Architectural Shifts to Cut Token Costs and Latency

Stop Stuffing Your Context Window: 6 Architectural Shifts to Cut Token Costs and Latency

Comments
3 min read
A 20% Security Tax Is the Most Honest Number in AI Right Now

A 20% Security Tax Is the Most Honest Number in AI Right Now

1
Comments
3 min read
CoSnitch Is a Reminder That Your Chatbot Will Tell on You If You Ask Nicely Enough

CoSnitch Is a Reminder That Your Chatbot Will Tell on You If You Ask Nicely Enough

Comments
3 min read
Enclave vs. the agent frameworks we audited — it's a layer map, not a fight

Enclave vs. the agent frameworks we audited — it's a layer map, not a fight

Comments
5 min read
LangChain vs LlamaIndex vs Chonkie: same 94-page PDF

LangChain vs LlamaIndex vs Chonkie: same 94-page PDF

Comments
7 min read
My Agent's Tests Were Green Because the Model Learned to Cheat

Comments explore fixing broken evaluators

My Agent's Tests Were Green Because the Model Learned to Cheat

24
Picked as gem Comments 21
5 min read
Why GitHub Access Isn't the Same as Local Repository Access

Why GitHub Access Isn't the Same as Local Repository Access

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.