close

DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost

I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost

Comments 1
6 min read
瀏覽器跑 LLM 的實際體感與骨感限制

瀏覽器跑 LLM 的實際體感與骨感限制

Comments
2 min read
Embeddings Cannot Say No: An Intent Detector's Real Numbers

Embeddings Cannot Say No: An Intent Detector's Real Numbers

BERJAYA 2
Comments 1
5 min read
Moderating Large-Volume User Content: Batch LLM Cost Estimates and Token Counting

Moderating Large-Volume User Content: Batch LLM Cost Estimates and Token Counting

Comments
7 min read
Batch LLM Jobs vs Realtime APIs — Bulk Summarization Cost Attribution

Batch LLM Jobs vs Realtime APIs — Bulk Summarization Cost Attribution

Comments
6 min read
The reward signal is the bottleneck, not the model

The reward signal is the bottleneck, not the model

Comments
3 min read
Free Coding Models Are Good Enough for Some of Your Tasks — Here's How to Find Which Ones

Free Coding Models Are Good Enough for Some of Your Tasks — Here's How to Find Which Ones

Comments
4 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Comments
7 min read
Let the Test Suite Decide Which Model Answers: A Verification-Gated Model Ladder

Let the Test Suite Decide Which Model Answers: A Verification-Gated Model Ladder

Comments
6 min read
Route by Task, Not by Hype: A Budget-Aware Harness for Trying New Coding Models

Route by Task, Not by Hype: A Budget-Aware Harness for Trying New Coding Models

Comments
5 min read
Running three AI models on one local server when your VRAM doesn't cover all of them

Running three AI models on one local server when your VRAM doesn't cover all of them

BERJAYA BERJAYA 3
Comments
3 min read
How LLMs actually call functions (they don't)

How LLMs actually call functions (they don't)

Comments
5 min read
Don't Start With RAG: Lessons From Building an Automotive AI Pipeline

Don't Start With RAG: Lessons From Building an Automotive AI Pipeline

BERJAYA 1
Comments 1
5 min read
A Sandbox Got Popped at Black Hat. Nobody Should Be Shocked.

A Sandbox Got Popped at Black Hat. Nobody Should Be Shocked.

BERJAYA 1
Comments
3 min read
Red-teaming your AI agent like an attacker would: introducing Argus

Red-teaming your AI agent like an attacker would: introducing Argus

Comments
3 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.