English · 中文

Evaluate RAG Before Tuning It: A Minimal Baseline for Retrieval, Citations, and Abstention

Mar 14, 2026

Use a small deterministic data set to separate retrieval recall, citation support, grounded answers, and abstention instead of replacing evidence with one LLM-judge score.

Building a Traceable Document Pipeline with DeepSeek-OCR 2

Mar 3, 2026

Build provenance around DeepSeek-OCR 2 from immutable source and page rendering through pinned model execution, structural parsing, and human review.

A Claude Sonnet 4.6 Computer-Use Tutorial That Starts with Safety Boundaries

Feb 21, 2026

Place Sonnet 4.6 computer use inside an isolated environment with domain and action allowlists, action-time approval, prompt-injection resistance, and replayable audit.

Claude Opus 4.6 for Long-Context Agentic Coding: What Matters

Feb 10, 2026

Separate Anthropic's published Opus 4.6 capabilities from local evidence, then evaluate retrieval, editing, verification, effort, and total task cost.

Mistral Vibe 2.0: Do Custom Subagents Make Terminal Coding More Reliable?

Jan 31, 2026

Evaluate Vibe 2.0 through subagent contracts, clarification, skills, permissions, and repeatable repository tasks instead of counting product features.

  • « Prev
  • Page 6 of 14
  • Next »
Portrait of Lukes Lu

Lukes Lu

Software Developer

  • lukes.lu@yahoo.com
  • github.com/ilukes
  • @lupingui
  • Home
  • About
  • Keywords →
    • SwiftUI
    • OpenAI
    • Responses API
    • cancellation
    • iOS 26 beta
    • reasoning effort
    • AI Agent
    • Codex

©2026 All rights reserved. Made with Jekyll and ♥