Tech Big Bang
Daily digest of yesterday's top tech news and highlights
10+ years in automotive systems, AI application explorer, indie developer. Sharing hands-on experience in in-vehicle development, AI engineering, and team management, while incubating small but beautiful products.
Technical articles, tutorials, and insights covering automotive dev, AI engineering, and team management
Three embedding strategies tested on 27 Python functions: Strategy A raw code (Recall@5=0.958), Strategy B signature+docstring only (0.917), Strategy C hybrid (0.917). Counterintuitive finding: raw code embedding performs best. Root cause analysis: technical vocabulary in code (library names, exception classes, parameter names) provides high-quality semantic signals that docstring-only strips away. A second finding: some queries can't achieve perfect recall under any strategy, revealing genuine semantic gaps.
DataElem's open-source enterprise LLM application DevOps platform, named after Bi Sheng (inventor of movable type). Integrates AGL-powered Linsight Agent, visual workflow builder with Human-in-the-Loop, high-precision document parsing (trained on 5+ years of proprietary data), enterprise RAG, and unified model management. 11.6k Stars, Apache-2.0.
Local-first web intelligence layer for AI agents. 10 tools: search, fetch, crawl, extract, cache, find_similar, research, agent, diff, watch. 18 direct search engine adapters, ML reranking, byte-pinned source provenance, local semantic cache, no API keys required, $0/query. Comparison with Tavily, Exa, Firecrawl. 1.2k+ Stars, AGPL-3.0.
Local-first code intelligence graph. Tree-sitter parses the AST, SQLite persists the graph, 30 MCP tools surface precise context. Blast-radius analysis traces change impact chains, incremental updates complete in 2 seconds, 40+ languages supported. Benchmarked 82x median token reduction (528x max). 19.7k+ Stars, MIT.
A codebase isn't an upgraded document store β it's a fundamentally different knowledge type. Four understanding layers (syntactic/semantic/architectural/intent) map to four retrieval requirements. Traditional grep, vector search, AST symbol indexing, and call graph retrieval each have distinct capability limits. This article builds the complete framework for codebase knowledge systems.
A framework comparison of six codebase knowledge tools: codebase-memory-mcp, Cursor Context, GitHub Copilot Workspace, sourcegraph/zoekt, Tree-sitter, and OpenHands CodeBrowser. Five standardized test tasks (symbol location, semantic search, impact analysis, architecture understanding, history tracing) reveal each tool's capability boundary. Ends with a scenario-driven selection framework.
Independent products, from idea to launch β all by myself
AI skill marketplace β build your personal AI avatar and enterprise digital employees
Tell your travel stories with maps, record every journey
Goal-driven task management that turns plans into action
Your personal AI assistant that knows you better every day
AI-powered smart notes for effortless knowledge management
R&D quality management tools for continuous process improvement
Sharing tech insights and life thoughts on podcasts
Daily digest of yesterday's top tech news and highlights
A knowledge-sharing podcast focused on health and longevity β helping listeners discover cutting-edge health insights and practical strategies for a longer life
Digging beneath the surface of trending social events to uncover the stories behind the headlines
Curated content delivered to your inbox weekly β tech insights, product updates, and industry trends. No spam, unsubscribe anytime.
We only collect your email for the newsletter and will never share it with third parties. Unsubscribe anytime.
1,024 subscribers Β· Privacy-first, never shared
Find Me
Follow me on these platforms for the latest updates