Work is changing. Your space is how you answer.
Following that change — through real workspaces, live signals, and a yearly festival.
WDSF 2026 — winners on record
Off-season · See the 2026 resultsDesk Setup of the YearChen Sifan · Hangzhou, China · S·0068
Why this ledger exists
- 01
Judgment and taste
When AI does more of the doing, the human part of work gets sharper — judgment, taste, direction.
- 02
An observation post
AI is rewriting the workday in real time. Beyond Desk watches where that change becomes physical.
- 03
Authorship
How you work is becoming something you design, not something you are given.
Setups. 178 real workspaces, logged as submitted — people first, gear second.
browse all →
S·0057Student遇到困难睡大觉 · Sydney, Australia
S·0121Software developerMars Xiang · Shanghai, China
S·0051Flight attendant一只狗 · Shanghai, China
S·0052Abdullah bin Mohammad · Saudi Arabia
S·0053PhotographerAkayu · Guangzhou, China
S·0054EducatorAlana (The Simple Norm) · United States
S·0055Software EngineerAlex · Boca Raton, FL
S·0056Product designerAlex Richard · Dongguan, China
Scan. The one tool here that reads your own desk — a private AI report, by email.
scan your desk →One photo, read as a working system — scored 1–5 on the same four dimensions the festival uses, each with a written reason.
- ReportArrives by email — no total, no ranking
- PrivacyPrivate to you — kept out of the gallery and the registry
- StorageHeld in access-controlled storage, then deleted on a fixed schedule
The four dimensions it reads
- 01Work-mode fitDoes the layout serve a believable, describable workflow?
- 02Spatial narrativeCan a stranger read the person and place from the frame?
- 03Craft & executionHow completely is the intent finished — not how much it cost?
- 04AuthenticityA space in real use, or staged for show?
Blog. Thinking built on the evidence — essays and notes reasoned from recorded signals.
all entries →- B·0005
Safe models do not compose into safe systems
Six papers in two weeks converge on one engineering fact: safety measured on a single agent does not survive wiring agents together. The failure lives in the joints — memory, context, setup files, and the management layer itself.
- B·0006
Delegation regret is measurable now
Three studies put structure on a feeling agent users know well. The precise version: people regret the scope an agent took, not the errors it made.
- B·0004
This week in the record: agents are leaving the IDE
Six recorded entries from one fortnight point the same direction: the coding agent's home is no longer the editor buffer. It is Slack, Linear, the issue tracker, the phone.
Evidence graph
The ledger as a map. Dashed edges are machine-suggested (embedding similarity and duplicate clusters); solid edges are editorial — they appear only where a blog post cites an entry.
- cites — editorial citation (blog post → entry)
- related — machine-suggested (embedding similarity)
- cluster — same story, archived duplicate
AI hardware & peripherals
- Adjust your Navigator Trackball angle and sensitivity
- Press to exit auto-mouse and other Navigator firmware updates
- iFLYTEK P1 Pro: The AI-powered recorder slash lighter
- DIY Voyager Angled Keycaps
- Layout Buffet - Backlighting
- INNOCN CB32U1 review: An affordable alternative to the Apple Studio Display?
- Flagship convertible with Snapdragon X2 Elite - Microsoft Surface Pro OLED 2026 Review
- Excellent business laptop with 64 GB RAM - Lenovo ThinkPad T14s Gen 7 AMD Review
- June Roundup of Navigator Updates
- Layout Buffet - Hyper, Meh, and Software Shortcuts
- DIY Navigator Moonlander Shells
- DIY Moonlander Wrist Rest Pad
- DIY Navigator Thumb Module
- DIY Navigator Trackpad Shells
- iFLYTEK AI Recorder P1: An AI assistant as a wearable device
- AI at the Edge is a different operating environment
- HP Reveals Keyboard Computer with Ryzen AI Chip
- Keychron's Nape Pro turns your keyboard into a laptop‑style trackball rig
- Exploring Different Keyboard Sensing Technologies
- Toucan Wireless Split Keyboard with Touchpad
- OpenAI and Broadcom unveil LLM-optimized inference chip
- 192GB of VRAM in One PC… The Cheap Way
- AMD Built the DGX Spark Rival I Predicted… But There's a Catch
- Building the World’s Smallest AI Workstation
- Most Powerful 16" Gaming Laptop Has a Secret
- NVIDIA RTX Spark Made Everyone Mad
- 3 New PCs, One Giant AI Model… This Shouldn’t Work
- Three months wrong about why my 4-node AMD cluster was slow
- The Non-NVIDIA AI Card Everyone’s Ignoring
- INSANE $1000 32GB VRAM Local Ai Server
- $1500 Local AI Server Build Tested with Hermes Agent Gemma 4 and Qwen 3.6
- This $129 AI Storage Server Upgrade Saved Me $2,800
- BEST Local Ai Motherboard CPU Combo You've Never Heard Of
- KEEP AI LOCAL! Explaining Agentic AI and The Loop: FT MSI Cubi NUC+ and the MSI EdgeXpert Mini PCs
- Is STRIX Better than SPARK? Now Launching w/new Software: AMD's Ryzen AI Halo Developer Workstation
- NVIDIA's Secret Windows Upgrades for N1/N1X Laptops Help Everyone; AMD Could Benefit the Most
- Minisforum's N5 Max is an Absurd NAS
- Analyzing Nvidia GB10's GPU
- Inside Nvidia GB10’s Memory Subsystem, from the CPU Side
- Nvidia’s B200: Keeping the CUDA Juggernaut Rolling ft. Verda (formerly DataCrunch)
- Today was the perfect day for Poolside to drop Laguna S 2.1 because I just got these in! Finally have a half decent amount of VRAM. 3x V620 = 96 GB.
- RTX 5090 and 5080 Ran the Same Local AI Until the VRAM Ran Out
- Got these baddies in the mail today (2X 3080 20GB)
- Apple’s Hidden AI Model… The Speed they never showed
- Apple M5 isn't making full use of its matmul cores yet
- How much are RTX PRO 6000s going for in your country/state?
- PSA: DO NOT use Intel consumer platforms for multi-GPU setups
- Building the World’s Smallest AI Workstation
- Taking a Look at Gigabyte's MW94-RP0 Xeon 6 E-ATX Motherboard
- Is it worth getting 128GB MacBook Pro? Will it ever be comparable to today’s frontier models for coding?
- AMD Says 2 Ryzen AI Halos Can Run a 400B Model... I Tested It
- World's First(?) Underwhelming AMD Ryzen AI Halo Cluster
- AMD Strix Halo with 32 GB RAM for $1899 - Asus ProArt PX13 2026 Convertible Review
- A user has managed to run Kimi K3 on 80xRTX 5090, via 25GbE Ethernet.
- My second Inspur AGX-2 with another x8 v100 arrived!
- PCIe Gen6 and Gen5 Will Both Matter for AI Storage
- DGX Station running GLM5.2 in VSCode
- ASUS Showcases NUC 16 Family Powered By Panther Lake
- "Data center in a Box (on Wheels)" 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks
- Checking Out The GMKTek X3 Strix Halo: More Strix Halo Shenanigans!
- SK hynix, In Collaboration With SanDisk, Unveils The New High Bandwidth Flash (HBF) Standard, Helping To Resolve AI Inference Bottlenecks, Targeting Up To 3TB/s Bandwidth
- Kimi K3 full model running on 16x GB10 cluster at 20+tps
- Powerful creator tablet with Snapdragon X2 Elite - Asus ProArt PZ14 Review
- Get AI max+ 395 laptop or wait for rtx spark?
- Google DeepMind reshuffle 🧠, Meta Muse Code 💻, Anthropic chip team 🧩
- We Upgraded Our ZOTAC 4090 24G to Have More VRAM!
- RTX 5090 Owner Built An Open-Source Tool That Shuts Down PC If It Detects The 12VHPWR Cable Drawing Too Much Power, But It Can Only Work On Specific GPUs
- A 10M IOPS Kioxia GP1 SSD Shown Running at FMS 2026
- Showoff Saturday: Local 4x 6000 Pro (multi-year progression)
- enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think
- RTX 5090 96GB spotted on Alibaba?
- Underestimated budget solution: radeon 780m iGPU
- WisdPi WP-UT9 USB 10GbE Adapter Review
- Comu Action Pro: AI Assistant put to the test
- Comu Action Pro: AI assistant put to the test
- Show HN: A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)
- Minisforum N5 Max Review with AMD Ryzen AI Max+ 395
- I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti
- Ling-3.0-flash quant ladder on one DGX Spark: the whole thing sits in a 32 to 40 tok/s band
- Low Power AI Is More Efficient than NVIDIA // The AI Hardware Show S2E10
- OdinLake O3 WireControl ergonomic chair review: Better than the LiberNovo Omni?
- Nvidia doubles RTX PRO 6000 Blackwell's MSRP to a staggering $16,000 — 96GB card started pre-orders below $8,000 last year
- You could purchase a Desktop with 2TB of DDR5 - It only sets you back some $200k+
- 27-inch monitor with a 240 Hz IPS panel, G-Sync, and accurate colors: KTC H27E6S review
- M5 MacBook Air Against Every Generation for Dev Work
- This was a data center a year ago… Now it's on my desk
- This 13-in-1 USB-C docking station with built-in GaN+ charger changed my charging habits for the better — a review
- club-5060ti refresh: tested RTX 5060 Ti presets, a proper high-context harness, and Qwen3.8 27B
- MOKiN's 13-in-1 USB-C docking station has an actually useful gimmick and it turned my Lenovo Legion Go into a usable desktop PC
- Show-off Saturday: Intel Arc B140 build.
- The dream is to reach 200GB VRAM
- Qwen3.8-27B Q8_0 on Strix Halo is seriously impressive
- CDW has bumped the MSRP of the RTX Pro 6000 from $16,000 to $19,999
- Linux Improves VRAM Management in 7.3 Kernel 🥳
- This laptop with 64 GB RAM is perfect for AI agents: HP ZBook 8 G2a 14 review
- Alibaba's RISC-V CPU, XuanTie C950, Runs Qwen-3.8 27B at 30 tps
- Framework 13 Pro...Why Apple Solders Memory
- A Server with EIGHT GPUs? ft. ASRock R9600D
- Buying a V100/older NVIDIA GPU? Run this to check for older memory issues
- “The All Spark” Cluster: Upgrading from 16 - 36 DGX Sparks
- GMKtec is going to launch new hardware with Ryzen AI Max+ PRO 495 at IFA Berlin 2026
- With AI lumbar support, massage, heating, and cooling: Hbada X7 office chair review
- Apple M5 Server
- Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory
- Apple releases M5 ultra at 1.2TB/s bandwith
- Intel Arc Pro B60 Dual 48G spotted
- Mac Studio M5 Max Cost Analysis
- Mac Studio M5 Ultra is a DGX Spark Killer for Local AI
- 224GB of GPU Memory on 1 Desk and It Should Not Work
- EliteBook X Flip G2i 14 AI review: HP's strongest convertible yet is also one of the priciest
- [AINews] Hot Chips: OpenAI’s Jalapeño, Cerebras CS-5, Groq 3 LPX, Apple M6
- Let's Talk about SR-IOV and Proxmox on Intel Arc B60 and B70!
- I Built a Monster PC Just to Run AI Locally
- ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI
- Hot Chips 2026: Samsung’s Processing-in-Memory (PIM)
- Exo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clustering
- DGX Spark cluster
- It's official! 192GB Framework
- Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max — speed vs context depth, 100 turns, one graph
- When you say, because I can. Limits of X870e
- Experience report - Qwen 3.8 Flash Next on memory rich, GPU poor setup
- NVIDIA® DGX Station™ Delivering Data-Center-Class Performance from the Desktop
- This $60,000 Mac Cluster Has a $10 Problem
- HP ZBook Ultra G1a 14 with Strix Halo review: A victim of the memory crisis?
- A very confusing report from Puget Systems
- HiDock P1 AI voice recorder hands-on review
- 2/5 of my CMP 170HX have died after 2 weeks and the 3rd came with defective tensor cores. Current prices DO NOT justify the risk you are taking
- GLM 5.3 Flash makes a black hole Minecraft mod running locally on 4x RTX PRO 6000 WS
- Could the shortage be getting better?
- Unholy Strix Machine! Doubling up with R9700s
- MINISFORUM MS-S1 MAX-P495
- GB10 Dgx Spark 2 Node Cluster From Scratch
- AMD unveils Threadripper Halo Station
- MSI EdgeXpert (DGX Spark) Unboxing
Agentic workflow patterns
- A couple of days ago, I sat down with Vivek Bharathi and dumped my brains. Here's the interview...
- a sneak preview behind an embedded software factory. I suspect rapid application dev is back
- porting software has been trivial for a while now. here’s how you do it.
- don’t waste your back pressure
- teleporting into the future and robbing yourself of retirement projects
- everything is a ralph loop
- i ran Claude in a loop for three months, and it created a genz programming language called cursed
- Note #728
- The twilight of the chatbots
- Claude Dispatch and the Power of Interfaces
- Three Years from GPT-3 to Gemini 3
- The Shape of AI: Jaggedness, Bottlenecks and Salients
- Management as AI superpower
- A Guide to Which AI to Use in the Agentic Era
- On Working with Wizards
- Real AI Agents and Real Work
- An Opinionated Guide to Using AI Right Now
- Why AI Infrastructure must evolve for Agent Experience — Akshat Bubna, Modal CTO
- [AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI
- Vercel's Andrew Qu on why agents are a new kind of software
- The website of the future may assemble itself for every visitor
- Skill engineering and the case against one-shot AI design
- How Cursor deploys AI inside the enterprise
- Warp CEO Zach Lloyd on why software factories are the next phase of coding
- Autoresearch: The feedback loop behind self-improving agents
- Forward Deployed Engineers and the future of software engineering
- AIEWF Daily Dispatch: Loops, Software Factories & Forward Deployed Engineers
- Co-Existence and the End of Co-Intelligence
- CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions
- Scaling Managed Agents: Decoupling the brain from the hands
- How we contain Claude across products
- Building a C compiler with a team of parallel Claudes
- How we built Claude Code auto mode: a safer way to skip permissions
- Harness design for long-running application development
- Demystifying evals for AI agents
- Effective harnesses for long-running agents
- Code execution with MCP: Building more efficient agents
- How we built our multi-agent research system
- Effective context engineering for AI agents
- Claude Code: Best practices for agentic coding
- Building effective agents
- The "think" tool: Enabling Claude to stop and think in complex tool use situations
- Tuning the harness, not the model: a Nemotron 3 Ultra playbook
- Your coding agent bill doubled. Here’s how to fix it.
- Why Model Neutrality Matters More Than Cloud Neutrality
- Improving Agents is a Data Mining Problem
- Wiki Memory
- How to Use RLMs in Deep Agents
- How Candidly Built State-Aware Agent Harnesses with LangSmith
- Introducing Dynamic Subagents in Deep Agents
- Building Durable AI Agents
- AIUC-1: Building trust in AI agents
- Zero Trust for AI Agents
- Humility in the Age of Agentic Coding
- Agentic Coding and the Economics of Open Source
- Post-Mortem of Anthropic's Claude Code Leak
- Inside an AI-Run Company
- Beyond chatbots: Agents that tackle your SOPs
- While loops with tool calls
- Mosaic: Runtime-Efficient Multi-Agent Embodied Planning
- When is Routing Meaningful? Diversity and Robustness in Language Model Societies
- Secret Scanner Agent: Extracting Secrets and Access Context from Unstructured Documents
- Shared Selective Persistent Memory for Agentic LLM Systems
- Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting
- Rlm-Workflow
- Show HN: Durable AI agents without the workflow engine
- Claude and Codex and Grok: my current workflow and its friction
- Show HN: Sigil – FIDO2 key-derived P2P remote desktop (agentic workflow retro)
- "Code Is Cheap. Show Me the Talk.": Lessons from Teaching and Managing AI Coding Tool Usage in a Visualization Course
- Intervenability as a Design Requirement for Autonomy and Oversight within Human-Centered AI
- Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control
- Motif: Discovering and Automating Personal Web Workflows
- U-Lens: Supporting User Uncertainty Management in Long-Form LLM Responses
- Memory-Conditioned Tool Calling for Camera-First Visual Agents
- Exploring Agentic Workflows for Generating High Quality Math Visual Aids
- Neutralizing Structural Inequality in the Nigerian FinTech Sector
- Distributed Agent System: Fault-Tolerant Collaboration Among Embodied Agents
- Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games
- An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory
- Multi-Agent LLMs Fail to Explore Each Other
- How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm
- Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction
- Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems
- Agentic Context Learning with Self-Discovered Specification
- Verification of Adaptive Agentic Controllers through Finite Rule Revision
- Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems
- Automated Textbook Auditing with Multi-Agent LLM Systems
- Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
- Can Agentic Trading Systems Pay for Their Own Intelligence?
- Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers?
- StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure
- When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
- 5 Trends That Defined AI Engineering at World’s Fair 2026
- Compos3D: Interactive Part-Based Composition for Creative Control in Generative 3D Models
- A\"ira: Rethinking AI Research Assistants for Interdisciplinary Science
- Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
- SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation
- Self-Regulated Reading with AI Support: An Eight-Week Study with Students
- ParaTutor: Coordinating Parent and Child Math Tutoring through Role Separated LLM Scaffolding
- MetaInfer: A Knowledge Only LLM Inference Engine Generator SKILL Toolbox
- Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations
- RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls
- Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration
- XScientist: A Git-Like Research Protocol for Long-Running Autonomous Scientific Discovery
- Agent Identity URI Scheme: Topology-Independent Naming and Capability-Based Discovery for Multi-Agent Systems
- Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
- Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
- Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
- Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation
- When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects
- Learning Latency-Aware Orchestration for Multi-Agent Systems
- Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
- Benefits and Limitations of Communication in Multi-Agent Reasoning
- MASPRM: Multi-Agent System Process Reward Model
- Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications
- Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents — cited by 1
- Beyond Interestingness: Semantic and Context-Aware Natural Language Query Recommendations for Visual Data Analysis
- Proving the ROI of agentic AI in financial services
- Latent Communication Between Language Model Agents: Channels, Alignment, and the Limits of Text
- Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation
- ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
- Orchestrating Power Grid Studies with Multi-Agent AI and MCP Servers
- AI Agents Do Not Fail Alone:The Context Fails First — cited by 1
- When Is Delegated Play Truthful? Within-Range Regret and the Trilemma of Aligned Delegation — cited by 1
- Towards an Intention Abstraction Layer for Autonomous Industrial Systems
- Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems — cited by 1
- StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
- Does Multi-Agent Debate Improve AI Feedback on Research Papers?
- Proposed spec to share SKILL and "loop"/"Workflow" by OCI registry
- Ask HN: Workflow Automation vs AI Agents?
- Let AI take over your API debugging workflow
- Dealing with increasingly complicated agents
- We've all done RAG, now what?
- How Cars24 scales conversations and builds faster with OpenAI
- How to manage AI investments in the agentic era
- How sales teams use ChatGPT Work
- How data science teams use ChatGPT Work
- Australian Payments Plus moves faster with ChatGPT and Codex
- How agents are transforming work
- Codex-maxxing for long-running work
- Automating the Rubber Stamp: What If an Agent Ran Your Deployment Gate?
- Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k)
- The Pulse: What can we learn from Bun’s rapid Rust rewrite with AI?
- Context engineering with Dex Horthy
- What is “loop engineering?”
- Impressions from visiting OpenAI, Anthropic, & Cursor
- The Pulse: a trend of trying to cut back on AI spend within eng departments?
- Ideas: slow down to speed up when working with AI agents
- What building Shippy taught us about building agents
- Model Routing Is Simple. Until It Isn’t.
- Shipping huggingface_hub every week with AI, open tools, and a human in the loop
- MosaicLeaks: Can your research agent keep a secret?
- Is it agentic enough? Benchmarking open models on your own tooling
- How an Agent Built a 3D Paris Gallery by Chaining Two Hugging Face Spaces
- I was giving my coding agent context the wrong way...
- How to build proactive agents & self-improving company (Fully explained)
- Ralph-loop 2.0? The real autonomous coder is coming...
- 3 New PCs, One Giant AI Model… This Shouldn’t Work
- Breakdowns for Human-Machine Creative Reflexivity
- Perceived AGI: Believability as Dimensional Completeness, Not Capability
- When Not to Automate: A Formal Protocol for Human Preservation in AI-Optimized Organizations
- HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization
- Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation — cited by 1
- CoWeaver: A Bi-directional, Learnable and Explainable Matching Engine for Mixed Human-Agent Science Collaboration
- The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure — cited by 1
- Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing
- A hierarchical memory architecture overcomes context limits in long-horizon multi-agent computational modeling
- Tmux + Fable = Cut 35% less token
- Building Governed Agents: A Framework for Cost, Control, and Compliance
- TaskArtisan: Designing Composable Generative Widgets for LLM-Assisted Analysis
- Sidekick: Designing Communication for Effective Multitasking with Computer Use Agents — cited by 1
- Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
- Can LLM Code Explanations Adapt to Diverse Problem-Solvers' Needs?
- Large Language Models in Architecture Studio: A Framework for Learning Outcomes
- Automated Hardware Validation Test Plan Generation for Large Scale AI Datacenter Platforms Using a Generative AI Multi-Agents Architecture
- Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
- Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool Orchestration
- RELIC: Revealed Principles for Learning Interpretable Composable Skills in Multi-Agent Planning
- Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
- Autonomous Discovery of Wireless Communications Algorithms
- Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking
- MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
- FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering
- O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
- OMAC: A Holistic Optimization Framework for LLM-Based Multi-Agent Collaboration
- FOCAL: Filtered On-device Continuous Activity Logging for Efficient Personal Desktop Summarization
- ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
- How ZigZag made our workflow worse before we fixed it
- Ask HN: I stopped fighting AI over-reliance and built a workflow around it
- Workflow time travel to prevent agent Vision Drift
- The Archaeologist’s Copilot
- DSLs Enable Reliable Use of LLMs
- Fragments: July 13
- Experiences with local models for coding
- Fragments: July 6
- Building Reliable Agentic AI Systems
- Fragments: June 16
- Fragments: June 2
- Fragments: May 27
- The test suite as a regression sensor
- The VibeSec Reckoning
- Bliki: Vibe Coding
- Fragments: May 14
- Bliki: Interrogatory LLM
- What is Code
- Fragments: May 5
- Fragments: April 29
- Structured-Prompt-Driven Development (SPDD)
- Fragments: April 21
- Feedback Flywheel
- Harness engineering for coding agent users
- Encoding Team Standards
- Do Automated Evals Work?
- “It’s Hard to Eval” Is a Product Smell
- The Revenge of the Data Scientist
- Why I Stopped Using nbdev
- FORGET Loop Engineering. Agentic Engineering is about THIS
- Claude Fable 5 BANNED: The First Model Agentic Engineers DON'T NEED
- Pi Coding Agent Observability: HTML Specs with Gemini 3.5 Flash and GPT Image 2
- Top #1 Opportunity for Senior Engineers: Agentic Engineering
- Pi to Pi: Two-Way Agent Orchestration with the Pi Coding Agent
- GPT-5.5 VERIFIED Opus 4.7: A Pi Coding Agent That REVIEWS Like YOU
- MAXIMIZE Your Claude Code Subscription (Without Getting BANNED)
- How Agents Manage Other Agents: Four Subagents Patterns in 2026
- How Autoresearch will change Small Language Models adoption
- Agents: Inner Loop vs Outer Loop
- Can We Close the Loop in 2026?
- The Agent Client Protocol Overview
- The importance of Agent Harness in 2026
- Context Engineering for AI Agents: Part 2
- Why (Senior) Engineers Struggle to Build AI Agents
- Agents 2.0: From Shallow Loops to Deep Agents
- The Rise of Subagents
- Pydantic AI 2.0: The New Best Way to Build AI Agents is Composing Capabilities
- The Best AI Coding Setup Isn't the Most Autonomous One (Here's Why)
- Finally, an Open Standard for the Karpathy LLM Wiki is HERE
- Google Just Dropped a Masterclass on Agentic Engineering (It's SO Good)
- The Creators of Claude Code and OpenClaw don't Prompt Their Agents Anymore?!
- Patterns for Building Cybersecurity Evals
- Using LLMs to Secure Source Code
- How to Work and Compound with AI
- Product Evals in Three Simple Steps
- Software Factories, Light and Dark
- Own the Outer Loop
- Earning taste and judgment
- Agentic Autonomy Levels
- The New Software Lifecycle
- Agentic Code Review
- Loop Engineering
- The Intent Debt
- The Orchestration Tax
- Understanding is the new bottleneck
- Code like a surgeon
- On owning a codebase, and why it may be the hardest job in software
- Automating Security Triage with HackerOne and Deep Search
- How we're using Sourcegraph and a Slack bot to detect vulnerabilities and react quickly
- Why coding agents fail in large codebases (and what to do about it)
- Building DataBot: Our always-on data assistant
- The Coming Loop
- SEPs Are Moving to Pull Requests
- The Coding Agent Is Dead
- Go Deep
- Tab, Tab, Dead
- Handoff, Please
- Stick a Fork in It, It's Done
- Software Is Made Between Commits
- Introducing Zed's Agent Metrics
- On Programming with Agents
- AI's 70% Problem
- MAIS: Exploring human-AI interaction in fair and transparent recruitment with a multi-agent LLM-based system
- Asymmetric Encounters with AI: Professional Designers’ Perceptions and Integration of AI Tools
- Crafting AI Explanations for Every Role in Your Enterprise
- Vibe Architects: Agentic Vibe Coders
- The Core Skill of Design in the AI Era: Critique
- Context Architecture
- Embracing AI with Claude's C Compiler
- Fragments: July 21
- How Apollo Rebuilt Its AI Assistant on Deep Agents to Power the Full GTM Loop
- Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent — cited by 1
- AInimation: Animating from Prompt to AI-Generated Responses
- EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
- HALO: Interactive Co-abductive Reasoning in Scientific Hypothesis Generation
- A Decision-Centered Reference Architecture for Trustworthy Agentic Commerce
- Engineering Trustworthy Agentic AI for Critical Systems
- Solve the CyberGym benchmark
- 3 Years of Graph Engineering with LangGraph
- Multiplayer
- Building an AI-orchestrated publishing workflow for a long-form writing project
- Agentic Workflow's Cache Keepalive Costs 8x Too Much
- Loop Engineering Workflow Based on Anthropic and Google Papers
- NTT DATA Group cuts incident analysis to 30 minutes with Codex
- How to Actually Run Your Coding Agent Safely (And Avoid the Horror Stories)
- A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace
- Not Birds of a Feather: Personality-Based Partner Selection in LLM Agents
- Knowledge-Centric Self-Improvement
- Show HN: Hanesu – An experimental workflow layer for AI coding agents
- Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
- Copilot cloud agent for Linear is now generally available — cited by 1
- Exploring the Design Space of LLM-Based Programming Support in CS Education: A Scoping Review through the Lens of Assistance Governance
- Thinkink: 2D Spatial Ink-native Interaction with LLMs
- pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development
- FedAgentKE: Federated Semantic Knowledge Evolution for Heterogeneous Agents
- ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems — cited by 1
- Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events
- Workload-Aware Caching for Multi-Agent Systems
- MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation
- UX-Context Design: Using UX Knowledge to Inform AI-Generated Design
- How much are you actually using your local models these days? Which ones do you reach for the most?
- What does it mean to "own your intelligence"?
- Deepseek V4 flash - Hy3 or is Qwen3.6 27B still the most solid for agentic/coding?
- Control panels to clarify user intent with Large Language Models
- Reliability-Contagion Feasibility in LLM Multi-Agent Networks
- When Language Models Meet NeuroGraphs: Exploring Enhanced Agentic LLM Framework Towards Brain Network Analysis
- Where FactsGo Missing: A LayerwiseTaxonomy and Per-Layer Attribution of Information Omissionin Air-Gapped LLM Agent Pipelines
- TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI
- My Ollama box picks the music now: an agentic DJ running on a 9B model
- Evaluating Agents Beyond the First Prompt
- Reflections and Recommendations on AI Adoption Practice from a Mixed-Ability Research Group
- The Help Ladder: Skill-Adaptive Peer Scaffolding for Real-Time Collaborative Programming
- From Vibe to Code -- and Back: Lexical Oscillation in the Formation of Design Intent with Generative AI
- Beyond Conversations: Spatially-Anchored Previews for Intent Disambiguation in LLM-Assisted Geometry Editing in Virtual Reality
- A Taxonomy of Confabulations and the Perception-Reality Gap in LLM-Assisted Immersive Scene Editing
- Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks
- A Comparative Study of MCP and A2A for Inter-Agent Coordination in LLM-Based Systems
- A Vocabulary for Multi-Agent Automated Research Systems
- How we built LangChain’s agent-first data stack
- The Orchestrator's Tax
- Compliance-First AI: Proving Agent Provenance for Regulated Engineering Teams
- Codex from 0 to 10M Users: Building ChatGPT Work — Akshay Nathan, OpenAI
- Show HN: Formally verified 3D CSG: Trust 93 lines spec, not 1000 lines AI code
- How building software is changing at Anthropic — cited by 1
- Show HN: Tines 3B – safe workflow automation for when everyone builds software
- Scientific computing in the age of agentic AI
- Stop Writing for Me: Generative Refusal in AI Tools for Thought
- Language as a Material Interface for Creative LLM Interaction
- What Gets Lost When Memory Becomes Media? Evaluating AI-Generated Oral History Visualization
- Agentic AI-enabled discovery across large-scale sleep physiology
- SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
- ARCHER: Agentic Rule and Compliance Harness for Executable Regulations
- CHILL-Harness: Counterfactual Harness Learning for Efficient Reasoning in Long-Horizon Agents
- Towards a Systems Foundation for Agentic Cloud Management
- Aethel: A Reproducible Graph-Retrieval Framework for Multi-Hop Financial Diligence
- Who Cares About the Model?
- Formal methods with Hillel Wayne
- One Run Is Not an Idea: The Implementation Lottery in Automated Research
- Living-Harness Is an Interactive-Agent Evolver
- DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution
- When Should AI Follow? Task Structure and Joint Adaptation by Human and AI Agents
- Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
- Ontologies Are So Back: Why AI Agents Are Reviving the Semantic Web
- Loop engineer practice #1: Reddit loop grew 0 to 95 Karma in 7 days
- The Economic Benefit of Refactoring
- Stacked pull requests are now in public preview
- Auditing Emergent LLM-Agent Collaboration through Cooperation-Obligation Coupling
- Scaling LLM-Driven Multi-Agent Systems: Design Principles and Architectural Scalability Analysis
- AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration
- Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale
- A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response
- LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
- Univé builds an AI-ready workforce
- Show HN: What should the GUI for AI agents look like?
- Inkling-Small 🧠, GPT-5.6 price cuts 💸, Gemini Robotics 2 🤖
- The Conductor Developer
- Evaluating code review agents with ReviewBench
- Are you ready for Le Chaton FAT or still wasting money on GPUs?
- Claude for ADHD: The Coding Workflow I Built for My Brain
- Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration
- TransMem: Transforming Hidden States into Memory for Large Language Models
- SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery
- Beyond Byzantine: An Organizational Consensus Algorithm for Self-Interested Agents Under Information Asymmetry
- Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search
- My Super Simple Software Factory (For Agentic Engineers)
- How Stripe Built their Knowledge AI Platform: A Company-Wide AI Agent on Deep Agents, Live in 1 Week
- Show HN: "Hedgehog" - An Opinionated workflow for building with AI
- ReVoicer: Conversational Voice Annotation for Human-Centered, LLM-Assisted Peer Review
- Revibing Code from Papers: Reimplementing HCI Artifacts
- Width, Memory, and Delay: A Resource Accounting for the Limits of Flat Multi-Agent Systems
- MAPLE-Guard: Memory-Aware Link Enforcement Against Memory-Link Poisoning in Multi-Agent Systems
- BANDMAS: Causality-Inspired Semantic Packet Scheduling for Bandwidth-Efficient Multi-Agent Collaboration
- HIERA: Hierarchical Multi-Agent Relevance Assessment for Content Discovery Systems
- Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution
- Training Small LLMs as Spatial Multi-Agent Policies
- Customize the reasoning level for Copilot cloud agent
- Trigger Copilot automations with comments
- Customer Experience (CX) Agents in Production: Lessons from Lyft, Vodafone, and LATAM Airlines
- Show HN: Gigacode – the model writes its own multi-agent workflow, then runs it
- How to evaluate voice agents: execution, outcomes, and experience
- Unpacking ChatGPT Work: the Agent for a Billion Users
- Stateful Governance for Concurrent Agentic Systems
- Emergence of Biased Consensus in Multi-Agent LLM Debates
- SABRE: A Multi-Agent Approach for Selecting Out-of-Distribution Detectors Under a Budget
- Group Perspective Matters: Regulating Debate Relationships Can Mitigate Blind Conformity in Multi-Agent Debate
- An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures
- Agent memory layers don't need an LLM deciding what to remember
- How we build an autonomous SRE Agent for Kubernetes Deployments
- Enacting Constructive Conflicts with AI Agents to Enhance Reconsideration among Novice Interaction Designers
- LEGOUI: Designing with UI-DSL Bricks to Balance Transparency and Controllability
- IntentLint: Supporting Intent Scaffolding and Prompt-time Linting in Human-AI Collaborative Data Analysis
- Preference-Driven Online Adaptation for Personalized Interaction Initiation in Proactive AI Assistants
- Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems
- Continuous Improvement and Parallel Autonomous Exploration: An LLM-Agent Framework for Searching Large Solution Spaces
- HELENA:Hierarchical Sparse Coordination over a Union of Complementary Topologies for MAS
- CURATE: Leveraging LLM Agents to Compose, Catalog, and Deploy Reproducible Workflows
- MIDAS: Multi-LLM Iterative Data-Adaptive Summarization
- Structured LLM Reasoning for Zero-Shot Human--Robot Coordination Under Hidden Goals
- Models, Harnesses, and Multi-Agent Systems
- Crazy AI workflow for banks/finserv – internal audit
- A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems
- Certifying Collective Reasoning in Multi-Agent Systems via Koopman Spectral Analysis
- ASGE-RR: Agentic Service Graph Embedding with Revisable Reservations for Dynamic AI-Agent Calls
- DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data
- Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Plan Coordination
- How HSP GRUPPE builds AI capabilities for tax advisory
- Your AI Second Brain Is Slowly Rotting (Here's How to Fix It)
- [AINews] Zawinski's Law of MultiAgents
- Speculative decoding in a tools call
- Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration
- Fact-Check Your Information (FYI): A Design Probe to Understand How People Actually Fact-Check Data-Driven Articles
- Strategy-first synthesis planning for complex natural products
- ADIAS: Automated Design of Interactive Agentic Systems
- Reflex: Demonstrate a GUI workflow once, replay it with zero LLM calls
- Model ML completes finance work more efficiently with GPT-5.6 Sol
- What building an AI-native finance function taught me
- The death of AI workflow builders
- How Zapier transformed core marketing processes with ChatGPT Work
- Virgin Atlantic sharpens customer journeys with ChatGPT Work
- From Human-Centered Design to Human-AI Collaboration: Why the Future of HCI Still Starts With People
- MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures
- Fluid Structure, Rigid Record: A Layered Organizational Design Framework for Agent-Native Organizations
- Muscle Memory for Agents: Compile not Merely Retrieve
- Beyond Tier Labels: Role- and Deployment-Dependent Model Substitution in Multi-Call LLM Workflows
- MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
- You Don't Need To Stay in The Loop: An Agentic Robotics Loop for Robot-Policy Improvement
- TDD inside the agent loop - theater or actual value?
- How many of your agent's calls actually need a frontier model?
- Thinking of ACE? We Can Do It with Fewer Tokens
- Show HN: Synapse – a monitoring SaaS shipped with my Claude Code workflow
- Building monday.com Sidekick: why capable agents need more than just tools
- Copilot memory and Ollama in GitHub Copilot for JetBrains
- Beyond Cash Flows: A Multi-Agent AI Framework for Valuing Clinical-Stage, Cross-Border Biotechnology
- ASCon: A Direction-Aware Reciprocal Agent--Step Contextualization Model for Failure Attribution in Multi-Agent Systems
- Reifying Research Logic: AI-Assisted Workflow Construction and Incremental Refinement for Quantitative Syntax
- Who Are You Explaining To? A Multi-Agent System for Audience-Aware XAI Narratives
- Persistent Recursive Worlds Enable Autonomous Software Evolution
- What is an AI agent?
- Stop being skeptical about AI for development with Charity Majors
- How RingCentral builds AI-native work from engineering to ops
- Socioduality: A Relational Process Framework for Human-AI Interaction
- When Do Institutions Beat Intelligence?
- Rethinking Agent Security as a Networking Problem
- Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
- MaSRead: Content-Addressed Reading of Replicated Latent Stores
- Harnessing agent memory to build lifelong AI partners for materials scientists
- Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier
- EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
- Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets
- What We Learned by Reproducing 2,200 papers from ICML
- Humans are Missing from AI Coding Agent Research
- Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles
- Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference
- Reconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, Point-in-Time Research
- Attune: A Self-Annotation Tool for Understanding Robot Operator Attention Profiles
- How to Build the Most Powerful System for AI Coding (Full Breakdown)
- React for Agents: Astro Creator Brings Hooks to his Meta-Harness, Flue
- Proxy-Validated LLM UX Micro-Simulations: An Artifact-First Protocol for Early-Stage Decision Support
- MobileMem: Learning from a Year of Mobile Experiences
- From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL
- A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents
- The Ultimate Guide to Making Your Entire Development Cycle AI Native
- FIXING Opus 5: PROOF that Prompt Engineering IS NOT DEAD
- AI Agents and the Future of VIS
- RaivenTracks: Branching Provenance for Conversational Visualization Workflows
- Who's Keeping Score? Interactive Steering of LLM-Powered Scoring with Attune
- BRA-Audit: Budgeted Runtime Auditing for LLM Multi-Agent Systems via Cumulative-Exposure Audit-Point Placement
- Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce
- ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems
- VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience
- The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines
- From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems
- A practical workflow for LLM-assisted development
- How Much Memory Does Your Agent Actually Need?
- Frontier Model Cost and Open-Weights Popularity is Driving Demand for Model Routing
- How NVIDIA scales expertise with ChatGPT Work
- Why This and Not That? A Collaborative Reflection Approach for Understanding Thought Coverage in Decision Making Support Dialog
- Appearing Legitimate is Not Enough: Interrogating Synthetic Agents in Representational Processes through a Participatory Design Lens
- Procedural Collapse: A Structural Account of Disengagement in LLM-Assisted Writing
- MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering model
- The Little Scientist: LLM Agent-Driven Discovery via the Scientific Method
- KernelArc: A Multi-Agent Framework for GPU Kernel Optimization
- Am I doing something wrong? Qwen 3.8 27B seems useless for agentic coding
- From Chrome DevTools to AI Engineering, with Addy Osmani
- Show HN: Grove, a formal workflow protocol for long-running AI coding agents
- Citizens Build, Agents Execute, Experts Govern
- LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents
- Measuring Proof Burden in Public Bounty Listings: A RentAHuman Case Study
- Report on The 1st Workshop on Human-Centered Proactive and Personalized Agents for Interactive Information Access at CHIIR 2026
- Contracting for LLM Delegation: Moral Hazard in Technology and Effort Choice
- A Locally Deployable Tool-Grounded LLM Multi-agent Framework for Automating Methane Emission Analysis and Reporting
- Bayesian Partner Modelling enables Adaptive Replanning for LLM Coordination
- Towards Reversible Forgetting: Managing Obsolete Knowledge in Continual Enterprise AI Agents
- CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
- DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning
- AME Agent Swarms Rewrite the Workflow
- The Pulse: We need to talk about migrations with AI
- AQuA's "self-improvement" updates research state, not the agent LM. What should a local port freeze?
- The /wayfinder Skill: Navigating the “Fog of War” of Planning
- I did it! I'm free! It's been 7 hours since I used claudecode
- Delegating or Doing? Understanding User Behavior in Hybrid Human-Agent Interfaces
- IRIS: Navigating and Reflecting on Writing Traces Using Intelligent Document Histories
- Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design
- Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses
- When Do LLM Agents Help? Deadline-Aware Mixed-Criticality Task Scheduling at the Autonomous-Vehicle Edge
- Show HN: A coding-agent workbench built around Matt Pocock's coding workflow
- Recursive Self-Improvement
- The new GitHub Copilot experience in Slack
- I tried to do agenic coding with Qwen 3.8 27B 3bit quant on a macbook air m2 24gb. It took 63 hours, but amazingly, the flight simulator worked.
- Show HN: I let an subagent workflow refactor my codebase for three days
- The Evolution of the Agent Harness
- Fast and Hard Code
- Adapting Fossil-scm as a platform for AI agentic workflow
- Has anyone actually made 64k feel like 300k+ with recursive local agents?
- 1/100 → 44/100: fine-tuning a 450M VLM on 50K browser screenshots
- Human judgment doesn't leave the software factory. It relocates.
- Practical Loop Engineering
- Agentic Code Quality
- The Belief Update Gate: Separating Inertia from Learning in Human-AI Interaction
- Edge-Based Agentic Retrieval-Augmented Generation for Autonomous FHWA Bridge Inspection Compliance
- Towards Traffic Modelling of Multi-Agent Systems: The Role of Coordination Topology
- Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Long-Horizon Workflows
- Intelligence EXPLOSION: Harness Engineering with Pi Agent, Deepseek, and Gemini
- Fragments: August 24
- How Toyota North America Put Enterprise AI on the Balance Sheet with Deep Agents and LangSmith
- Context Anchoring
- Humans and Agents in Software Engineering Loops
- Ask HN: Is there any way to use workflow to control Harness?
- Exploring Agentic Approaches for Data Issue Detection and Repair in AI-Assisted Visualization
- Multi-Agent Discovery and Resource-Aware Autonomous Exploration of Scientific Datasets
- Probing How Users Interact with Turn-Level Design Frictions for AI Chatbots
- "I want to be pushed, I want to grow": Enabling social workers to design evaluations of LLM augmentation in their work
- Opinion-Guided Layered Strategies for Decentralized Coordination
- PropUQ-MAS: Propagation-Aware Uncertainty Quantification for LLM Multi-Agent Systems
- Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep
- The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams
- Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models
- Today I merged the first feature branch written entirely by my 4060Ti 16GB!
- AI Proficiency: From Users to Builders
- How We Build Agent Environments & Tasks
- Building Self-Correcting Memory in OpenWiki
- Why Ramp built its own in-house coding agent, Inspect
- Markets, Not Planners: Decentralized Orchestration of LLM Agents with Private Information
- Agentopia on a Consumer GPU: A Reduced-Scale Long-Horizon Port with an 8B Model
- LLM Agents Perform Controlled Experiments Using Simulation Models
- Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
- MARS: Multi-Specialist LLM Relay System for Competitive Programming
- How loveholidays is making everyone a builder with Codex
- A minecraft clone I fully vibecoded with Qwen3.8-27b Q4
- Watch This If Your Coding Agent is Ignoring Your Rules (You Need Hooks)
- Agentic World Analysis (AWA) - an alternative way to explore systems and support decision making
- HypoForge: A Self-Improving Multi-Agent Framework for Automated Hypothesis Generation and Testing via Scientific Skill Learning
- Praxist: From Experimental Artifacts to Solution Lineages
- Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems
- Federation Is Nearly Free, Reasoning Is Not: Tradeoffs for AI Co-Scientists in Protein Characterization Workflows
- MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration
- Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory
- My software development workflow is AI now & it feels exhausting and soulless
- Making Your Data Ready for Agentic AI
- Show HN: Build your own theme park
- Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning
- Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance
- Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries
- One Model, Many Minds: Unlocking Multi-Agent Synergy in a Single Agent via Mixture of Roles
- Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy
- SKILL.state: Scalable Long-Horizon Agent Skills
- MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
- Building the Foundation for the Agentic AI Era
- Best workflow engine is a programming language
- Agentic Workflow Design: Six Principles for 2026
- AI Can't Replace Real Research in Empathy Mapping
- The Custodial Era of UX: Cleaning Up After AI
- The Hidden Flaw of EVERY Coding Agent Now Has a Solution
- Agency and Agents
- How Much Can AI Understand? Toward AI-Assisted Sensemaking of Collaborative Discussion in Groups with Shared History
- FocusGen: Expanding Visual Design Exploration with a Simulated Focus Group of Persona Agents
- AI as Teammate: Rethinking Task Distribution in Medical Training
- Between Algorithm (AI) and Intuition (Human): Preserving Designer Agency in AI-Assisted Sensemaking of Qualitative UX Data
- FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling
- Prove2Me: An Open Collaborative Platform for Scaling Math Formalization
- Offline-Verifiable Accountability for Cross-Organization Agent Messaging: A Preserved Evidence-Bundle Approach
- Logos: An Agent Harness on a Cross-Process Bus
- Space to Talk
- Agentic Engineering Operating Level: WHERE to FOCUS your AGENTS?
- GLM 5.3 and GLM 5.3 Flash ran locally on RTX PRO 6000 WS and built a penthouse using BlenderMCP
- Delegating Before Learning: Where Generative AI Sits in Students' Professional Communication
- Structured State Reconciliation for Human-AI Task Handover
- AREAs-Lab: An Interactive Environment for AI-driven Requirement Elicitation for AI Systems
- ASTRA - Agentic System for Ticket Resolution and Analysis
- Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps
- AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
- Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses
- Cognitive Cells: A Compositional Framework for Populations of Small Language Models
- How Foundational Models Became Superhuman in Bash
- PRs NOT Welcome: How Top AI Open Source Projects Are Managing Thousands of Contributors
- How AI-native companies turn workflows into operating capability
- Fragments: September 1
- Copilot code review can now approve pull requests
- Ask HN: What full workflow / process have you tried to automate with AI?
- How law firm Gilbert + Tobin governs and scales AI with OpenAI
- Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications
- Classic AI Scaffolding for LLM Social Agents
- Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
- EULER: Exploring Underused Links with Evidence-Checked Return for Multi-Agent Mathematical Discovery
- Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
- Qwen 3.8 Flash Next + HERMES AGENT = AWESOME LOCAL AI AGENTS!
- My Agentic Engineering Workflow after 6,775 sessions [video]
- Maybe We Shouldn't Be Reviewing All This Code
- An Accidental Blackboard
- ATV Big Air Tour turned 3 days of work into 3 hours with ChatGPT
- AI Software Factories Are the Next Big Thing (And I'm Building You One)
- Scaling Agents in Europe & The Middle East: Lessons from Schneider Electric, Vodafone, and monday.com
- Beyond Instruction-Driven Editing: Source-Grounded Problem Discovery with User-Governed Repair for Scientific Posters
- Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight
- OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations
- ArcticSwarm: Deferring Early Consensus in Long-Horizon Multi-Agent Research
- Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives
- Agents That Model Agents: Five Principles Toward a Theory of Mind for 6G Networks
- Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence
- Bonded Recourse for Smart-Contract Settlement of Compensable Agent Side Effects
- From Click-Ops to IaC: A Safer Workflow with AI
- Legora reviewed 41 documents in minutes with GPT-6 Astra
- GPT-6 Astra: an automated AI Engineer you can hire for <$6 an hour
- Exploratory Unstructured Data Analysis: A Formative Study and Implications for Human-AI Collaboration
- You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring
- Where Reliability Lives: Experimental Localisation of Behavioural Properties in an Agent System
- The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems
- The Illusion of Independent Quorums: Epistemic Fault Domains and Correlated Cognitive Failures in Agentic Quorums
- Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning
- Speculative Macro Commit for Faster Tool-Using Agents
- Using AI for UX Work: Study Guide
- I've found myself using Local LLM's like 3D printers.
- Qwen3.8 27b for agentic coding and next .... what?
Agent tools & setup
- Note #727
- Claude Code and What Comes Next
- Configurable AI Coding Assistants: Designing For Developers Who Like to Be in Control
- Introducing Zapp
- An update on recent Claude Code quality reports
- Quantifying infrastructure noise in agentic coding evals
- Beyond permission prompts: making Claude Code more secure and autonomous
- Introducing advanced tool use on the Claude Developer Platform
- Writing effective tools for agents — with agents
- Desktop Extensions: One-click MCP server installation for Claude Desktop
- Equipping agents for the real world with Agent Skills
- Introducing OpenWiki Brains, general-purpose wiki memory for agents
- Introducing Contextual Retrieval
- How Schneider Electric Built Their LLMOps Foundations At Enterprise Scale With LangSmith
- LangChain and NVIDIA launch the NemoClaw Deep Agents Blueprint
- Deep Agents Code on NemoClaw: a governed blueprint for your most sensitive code
- Harbor x LangChain: A Unified Stack for Evaluating Agents
- Introducing OpenWiki, an open source agent for repo documentation
- How Pendo used LangSmith to trace Novus from user behavior to code fixes
- Running Untrusted Agent Code Without a Sandbox
- Prompt Caching with Deep Agents
- Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet
- Rebooting Enterprise AI with MCP and Kubernetes
- Hermes Agent: Agents that grow with you
- Technical advances in document understanding
- Chris on AI, autonomous swarming, home automation and Rust!
- Beyond note-taking with Fireflies
- Show HN: Orchestrator – a single-binary workflow orchestration tool
- Ask HN: Would filesystem bookmarks be useful in your shell workflow?
- Bytechef open source platform for AI agent orchestration and workflow automation
- Show HN: DonnyClaude – a verified workflow engine for Claude Code
- Workflow State Engine – Ferricstore
- Show HN: Wayflow – an embeddable AI workflow builder (open source)
- Chainything: Workflow automation tool with no-code UI and AI assistant
- AI Specialists Ready to Transform Your Workflow
- Open-source AI agent workflow for auditing Solidity smart contracts
- Show HN: What if your menu bar was a keyboard-controlled command center?
- Your coding agents are a black box. Here's how to crack them open.
- Agents need their own computer. Here's how to give them one safely.
- New in LangSmith Fleet: Bring agents into Slack in one click
- AgentSociety 2: An Integrated Research Environment for Executable Social Science
- Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study
- SoftBoard: A Multi-Agent Tool for the Creation and Evaluation of Low-Fidelity Prototypes
- DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
- OpenWiki 0.2 is adopting the OKF support
- Unigent SDK – universal cross-harness, cross-session agent workflow scripting
- Show HN: Skillful, stop maintaining the same AI workflow in five places
- Show HN: AI Workflow Builder App Template for React
- Show HN: How you auto recover your Claude Code workflow when quota resumes
- Creating a private AI assistant in Thunderbird
- ChatGPT is now a partner for your most ambitious work
- Samsung Electronics brings ChatGPT and Codex to employees
- New usage analytics and updated spend controls for enterprises
- Show HN: Leaves – A text-UI disk usage treemap visualizer
- Show HN: Nobie – an Excel-compatible runtime for agents and humans
- Show HN: Clawk – Give coding agents a disposable Linux VM, not your laptop
- Show HN: Jacquard, a programming language for AI-written, human-reviewed code
- Show HN: Juggler – an open-source GUI coding agent, by the creator of JUCE
- The Pulse: Grok’s CLI caught uploading all your local files to the cloud
- Building OpenCode with Dax Raad
- Beta: Clawk
- Beta: dcg
- Beta: Mellea
- Thinking Machines Inkling 🧠, GPT-Red 🔒, Perplexity sandboxes 🛡️
- Claude Code browser 🌍, Cursor general agent 🤖, Claude Fable extension ⏳
- Devin Fusion 💻, DeepSeek DSpark ⚡, economy of tokens 💰
- Seedance 2.5 🎥, guide to Fable ✨, OpenAI preps GPT-5.6 🚀
- Profiling in PyTorch (Part 3): Attention is all you profile
- From Hugging Face to Amazon SageMaker Studio in one click
- Run AI workloads on any cloud, store on Hugging Face: zero-egress storage with SkyPilot
- 🤗 Kernels: Major Updates
- Run a vLLM Server on HF Jobs in One Command
- Experimenting with the proposed Cross-Origin Storage API in Transformers.js
- We got local models to triage the OpenClaw repo for FREE!*
- From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot
- Agentic Resource Discovery: Let agents search
- Migrating Your GitHub CI to Hugging Face Jobs
- The Open Source Community is backing OpenEnv for Agentic RL
- Designing the hf CLI as an agent-optimized way to work with the Hub
- wtf is Loop Engineer & how to setup for real
- New AI coding paradiagm - OpenAI Symphony
- Anthropic killed Tool calling
- WebMCP - Why is awesome & How to use it
- How to install and use Claude Code Agent Teams (Reverse-engineered)
- Your OS Changes Everything for Local AI
- This Is What Happens When You CRUSH An AI Video Model
- Visualization Autocomplete: Visualization Authoring via Stepwise Design Recommendations
- SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
- A Generative Partially Specified Finite State Machine Approach to Complex Behaviour Planning
- IssueBench - How We Evaluate Engine
- Octo-planner: On-device Language Model for Planner-Action Agents
- ETAS: An Effect-Typed Language for Agent Systems
- Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
- PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution
- Grabette: an open system to record robot-manipulation data
- The PERFECT Local AI Setup
- Hermes Agent the BEST Local Ai Agent
- Can YOU Run Deepseek V4 Locally?
- Viability of local models for coding
- Evals Skills for Coding Agents
- Selecting The Right AI Evals Tool
- Engineers... STOP Picking GPT-5.6 Sol OR Claude Fable 5… FUSE THEM
- SEE CMUX SOLVE Multi-Agent Orchestration (Claude Code and Pi Agent)
- PLANS For Fable 5: Rebuilding My /Plan Skill for Mythos Class Models
- Engineers, DELETE the BASH Tool: Agentic Security For Pi Agent and Claude Code
- My M5 Max, Gemma 4, MLX LOCAL Stack. (This KILLS MODEL PROVIDERS)
- Building Managed Agents That Use GitHub Without Exposing Your Token
- Control an Android Phone with Gemini 3.5 Flash Computer Use
- Getting started with the Gemini Interactions API
- How Gemini Managed Agents Works under the Hood
- Gemini Managed Agents: Developer Guide
- How to use Deep Research with the Gemini API
- How to correctly use MCP servers with your AI Agents
- 8 Tips for Writing Agent Skills
- Combine Built-in Tools and Function Calling in the Gemini Interactions API
- Practical Guide to Evaluating and Testing Agent Skills
- Writing a Good AGENTS.md
- Multimodal Function Calling with Gemini 3 and Interactions API
- Getting Started with Gemini Deep Research API
- Gemini Interactions API Quick Start
- MCP is Not the Problem, It's your Server: Best Practices for Building MCP Servers
- Building Agents with the Gemini Interactions API
- Introducing MCP CLI: A way to call MCP Servers Efficiently
- Practical Guide on how to build an Agent from scratch with Gemini 3
- Gemini API File Search: A Web Developer Tutorial
- Build your first AI Agent with Gemini, n8n and Google Cloud Run
- This Completely Changes the Way We Build Production AI Agents (Vercel Eve)
- I Turned Claude Code Into a Complete Video Generation System (with Archon)
- My AI Memory Now Follows Me Across Every Tool!
- Agentic Batch Changes is now in public beta
- Why your migration tools are failing your engineers
- Sourcegraph MCP server and a cheaper model beat a Mythos-class model alone
- Lessons on UX, security, and scale when building an enterprise-grade Slack agent
- Code Search, Deep Search, or MCP: When to Use Each
- MCP stories from the field
- A new era for Sourcegraph: The intelligence layer for AI coding agents and developers
- How our support engineers use Deep Search to investigate customer issues faster
- Why code search at scale is essential when you grow beyond one repository
- Fixing the React2Shell vulnerability in large and complex enterprise codebases (part 2)
- Omnigent: The New Meta-Harness for EVERY Coding Agent - Claude Code, Codex, Pi, More
- Google's Agents CLI: The CLI + Skills Combination to Ship AI Agents EASILY
- Better Models: Worse Tools
- Pushing Local Models With Focus And Polish
- Beta SDKs for the 2026-07-28 MCP Spec Release Candidate Are Here
- Enterprise-Managed Authorization: Zero-touch OAuth for MCP
- The 2026-07-28 MCP Specification Release Candidate
- Tool Annotations as Risk Vocabulary: What Hints Can and Can't Do
- Understanding MCP Extensions
- The 2026 MCP Roadmap
- MCP Apps - Bringing UI Capabilities To MCP Clients
- Exploring the Future of MCP Transports
- MCP joins the Agentic AI Foundation
- One Year of MCP: November 2025 Spec Release
- MCP Apps: Extending servers with interactive user interfaces
- Adopting the MCP Bundle format (.mcpb) for portable local servers
- Server Instructions: Giving LLMs a user manual for your server
- Update on the Next MCP Protocol Release
- Introducing the MCP Registry
- Meet Puck
- Amp Is Now In Slack — cited by 1
- Subscriptions, At Last
- From Agent to Agent
- Secrets of the Orb
- The Dial
- Agents, Anywhere
- More Orb Sizes
- Read Bigger Threads
- Agents in Orbs
- Custom Agents
- A Faster Librarian
- Diffs
- Faster Deep & Rush
- Agents, Everywhere
- Opus 4.8
- The End of Public Threads
- Plugins, Everywhere
- Drop the Neo
- Proof of Human
- GPT Image 2 Paints Better
- Rush, 2.0
- npm Package Changes
- Amp, Rebuilt
- GPT-5.5 In Deep
- Opus 4.7
- GPT‐5.4 in Deep
- GPT-5.4, The New Oracle
- GPT‐5.3‐Codex
- Liberating Code Review
- Slashing Custom Commands
- Painter
- Hidden Gems: Part 4
- What GitHub Copilot's Usage-Based Billing Means for Zed Users
- Terminal Threads Are Live in Zed
- Why and How to Run Local Models in Zed
- Use Your ChatGPT Subscription in Zed
- What Anthropic's New Claude Billing Means for Zed Users
- Introducing Zed for Business
- We're Not Building AI Features for the Money
- Introducing Parallel Agents in Zed
- How We Developed Zeta2
- We Rebuilt Zeta from the Training Data Up
- Choose Your Edit Prediction Provider
- The ACP Registry is Live
- Run Your Project in a Dev Container, in Zed
- Hidden Gems: Part 2
- Introducing Agent Extensions
- Codex is Live in Zed
- GitHub Code Quality is now generally available
- Repository-level GitHub Copilot usage metrics generally available — cited by 1
- Copilot code review: Customization and configurability improvements
- GitHub Mobile: Fix pull request comments with Copilot cloud agent
- Trace voice agents in LangSmith
- Gemini 3.6 Flash is now available in GitHub Copilot
- pi 0.81.0 adds support for llama.cpp
- Show HN: SciStudio: organize chaos data analysis pipelines with workflow runtime
- Right on Schedule
- Copilot users can now see AI credits used per billing cycle
- I ran Laguna-S-2.1 through my private agentic eval vs Qwen3.5-122B on an RTX Pro 6000 (96GB). Fastest 100B+ I've tested and the best tool calling, but it invents facts under pressure.
- AI Tool Discovery at Scale: All You Need is DNS
- Show HN: CodeAlmanac – Karpathy-style codebase wiki from your conversations
- Unsloth Quantization of Laguna S 2.1 Is Out
- Gigatoken: A new open source tokenizer ~100x faster than Tiktoken, -500-1000x faster than Huggingface
- Introducing OpenAI Presence
- Llama.cpp just added support for Laguna XS.2 & M.1
- Towards Automating Eval Engineering
- Show HN: Ipek – a visual IDE for workflow automations
- Show HN: Bento - An entire PowerPoint in one HTML file (edit+view+data+collab)
- New Copilot usage metrics impact dashboard
- MindControl - llama.cpp fork to guide the reasoning process via injection during sampling
- Tool: Databasement
- Beta: CodeAlmanac
- Beta: termcn
- Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong
- How We Benchmark Deep Agents
- Code Finder: fast, efficient code search for coding agents
- Event Driven Orbs
- PSA on Laguna S-2.1 - Use the updated chat template and GGUF
- July 2026: LangChain Newsletter
- Show HN: OneCLI – OSS credential gateway that keeps secrets out of AI agents
- Show HN: Palmier Pro – Open-source macOS video editor built for AI
- Show HN: Remux – an open-source tmux workspace designed for iPhone
- Show HN: DeepSQL – A self-hostable DBA agent for Postgres and MySQL
- GitHub Mobile: Fix failing Actions checks with Copilot cloud agent — cited by 1
- Agent automation controls in GitHub Issues in public preview — cited by 1
- GitHub MCP Server supports the next MCP specification
- How Codex became a collaborator for OpenAI’s creative team
- I compared local models and different quants / config on a subset of swe-verified bench
- [audio.cpp] Release 0.4: Higgs Audio v3 TTS 4B (10x real time)+ Fish Audio S2 Pro in C++/GGML, full GGUF loading, Q8 speed and VRAM gains
- HARP: The Human--AI Research Platform
- AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
- UPDATE - HuggingHack Is Now On Github
- Using the Bonsai 27b 1b quant locally - regularly.
- Demios – AI-native workspace for data capture, workflow automation and reporting
- CachyLLama’s: llama.cpp fork with persistent KV cache that makes long local-agent sessions much less painful
- DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)
- Python Toolkit: a GUI to manage python, venv, packages, reqs, AI interfaces and more...
- CachyLLama: llama.cpp fork with persistent SSD-backed KV caching for local agent workflows
- Show HN: Claude-thermos keeps your Claude session warm for you
- Llama.cpp now has full MCP support!
- POCKET-35B agentic model on cpu 59 t/s
- Minimax M3 support with MSA has been merged into llama.cpp
- Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash
- Enterprise managed settings in the GitHub Copilot app and Copilot cloud agent
- Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.
- Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure
- CRAFT: Learn the Schema, Execute the Plan
- Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline
- GitHub Copilot for JetBrains adds improved OpenTelemetry configuration and model management
- Show HN: Yap – OSS on-device voice dictation for macOS with no model to download
- Show HN: Whetuu – a zero-config cross-shell prompt written in Zig
- Show HN: I left VSCode to build an IDE to handle many projects/agents workflow
- GitHub Actions holds potentially malicious workflows for approval
- spec: add DSpark speculative decoding by wjinxu · Pull Request #25173 · ggml-org/llama.cpp
- Gemini Distillation Service
- DeepSeek V4 Flash, up to 32 tok/s on AMD Ryzen AI MAX+ 395
- The 2026-07-28 Specification
- Grok 4.5 is now available in GitHub Copilot
- I got Kimi-k3 running.....
- The Anthropic Economic Index connector
- Your OS Changes Everything for Local AI
- I tried running a 1.56TB MoE model on a 6GB RTX 4050 Laptop, Here’s the result
- Appliedin: The agentic workflow for applying jobs, so we can spend time prepping
- I built a GBNF grammar compiler that makes 8B models reliably call tools - here's how it works (deep dive)
- AI slowdown pact ⏸️, Personal superintelligence access 🌍, Grok Build Mode 🛠️
- The idea: on a CPU the decode speed depends on the active params per token, not the total. My objective is trying to run a 10B at 100tok/s on a mid level PC (No GPU).
- A slide deck you can edit with a local model or in Chrome — the whole deck is a JSON block in one HTML file (~640KB with editor and viewer included)
- How Similarweb Evaluates Long-Form Agent Research Reports with LangSmith
- From keyword to published post in one focused workflow
- Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
- Show HN: Bullshit Detector – agent skills that fact-check videos and articles
- Deep Agents v0.7
- Everyone posts day-one impressions. What's still in your stack a month later?
- Kimi K3 for local use (1.56TB → 594GB) compressed and released by Unsloth
- PSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled
- Ilintar's Official Guide To Model Selection
- Show HN: Qwen Scribe – local transcription and dictation for Apple Silicon
- Copilot code review: Agent skills and MCP now generally available
- Tool: superfile
- Beta: OpenWiki
- Beta: jcode
- The Ultimate Knowledge Base: Bring YouTube Into Your AI Second Brain
- Bought a 5090 to escape API fees. Ended up building a mini datacenter. Sound familiar?
- Anyone tested the IQ1_M 342GB Pruned Kimi K3? Is it usable?
- Benchmarked: MindControl for Llama.cpp
- Turbo-fieldfare: Open-source engine running Gemma 4 26B in 2 GB RAM on Apple Silicon
- GitHub Copilot in Visual Studio — July update
- LangSmith LLM Gateway: runtime controls for production agents
- Reference same-repository actions with self-repository syntax
- Limit remote control to managed devices
- GitHub Copilot in Visual Studio Code, July 2026 releases
- Show HN: Optimize and serve models with Fable quality at half the cost
- How avatarin built a 24/7 retail agent with GPT-Realtime
- VizPilot: Automated Onboarding for SVG-based Composite Visualizations using Multimodal LLMs
- Argonaut: Interactive Visual Exploration for Distributed Optimization
- VISA: A Structured Description Protocol for Agent-Based Simulation Models Towards Machine Reproducibility
- The Complete Local AI System with A Single NPM Install!
- Enterprise teams model policy targeting in public preview
- Show HN: Claude-account – switch Claude Code accounts without logging in again
- How to evaluate Sourcegraph on your own codebase
- [audio.cpp] Release 0.5: DramaBox expressive TTS, Confucius4 cross-lingual voice transfer, plus 7 more models and ROCm/HIP
- Weight-Aware Streaming Tensor Engine: run Kimi K3 using 29 GB of RAM at 0.50 tok/s
- DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.
- Fix for Deep Seek v4 Flash 0731 tool calling has been added to llama cpp
- DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5
- Koboldcpp v1.118 released
- Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG
- I pushed Kimi K3 onto one CPU with 8 GB of RAM
- DeepSeek-V4-Flash 284B on 5.3GB of memory
- Setting up of a 16xGB10 (DGX Spark) cluster
- llama.cpp just added MTP / DSpark support for DeepSeek V4 Flash
- Deepseek-V4-Flash-0731 Dwarfstar on Mac
- Show HN: Sprocket – The Best AI Agent for Hardware and Software Development
- Show HN: NixOS-DGX-Spark – Nix and NixOS on the DGX Spark
- CyberNeuro: A Privacy-Preserving Agentic Workbench for Cohort-Scale Neuroimage and Clinical Data Analysis
- The AnyLog Edge Data Fabric
- Show HN: Mu – Tools for Agents
- I CANNOT believe I've got DeepSeek-V4-Flash-0731, a frontier model, running on my home PC. Insane!
- DeepSeek V4-Flash (284B MoE) at 33 tok/s single / 68 tok/s aggregate on 2× RTX 3090 + a used quad-Xeon DDR4 server — full config
- Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
- can someone create a website where people share specific hardware specs with specific llama cpp flags so we see what works?
- Attach Anything
- Time to finally migrate from LM Studio -> llama.cpp, your experience?
- Sharing a persistent browser QA workflow for OpenCode
- Code coverage automatic enablement in Code Quality settings
- Deploy local agents everywhere with LFM2.5-2.6B
- Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
- [Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]
- Gemma 4 on 500MB
- Has anyone tried Mach-1 Additive? 95% of performance of Qwen 3.6 35B while being 10x smaller
- A llama.cpp PR caches “hot” MoE experts on the GPU — 33 → 56 tok/s reported with 8GB VRAM
- A 2.6B model with tool calling and 128K context now runs at 30 tok/s on a phone
- [AINews] Megakernels are so dead and so back
- MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
- Qwen3-TTS voice cloning is now in mainline llama.cpp — the old demo finally became real support
- Google LLM router ➡️, Cloudflare Wallets 💳, Anthropic and Volta 🤝
- Sandboxing
- Could we have a --disk-moe or --n-disk-moe like --cpu-moe or --n-cpu-moe so we can use disk/cpu/gpu ?
- Tool: Mu
- Beta: @cloudflare/computer
- Beta: key-amnesia
- The Creator of Claude Code Said to Do What Now?!
- Prime Agent - a new coding harness surpassing Codex/CC/PI
- EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis
- Portals into Orbs
- i just spent weeks rewriting my webUI from scratch, getting rid of all AI slop within the codebase and switching it over to a proper lightweight framework (alpine.js). i am now comfortable suggesting it as an alternative to openwebUI, librechat and the like! it is made for local models
- Auto-fit vs tuned MoE offload: 564 → 1330 pp tok/s, unchanged decode (Qwen3.6-35B-A3B Q6 / RTX 3090)
- Best llama cpp flags to run Deepseek-flash 0731
- I compared even more parsers on 14 PDF-parsing capabilities using different types
- nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging Face
- I ported vLLM's serving stack to C++20: 66 MiB binary, no Python at inference, output checked token-for-token against vLLM
- Deep Agents vs LangChain vs LangGraph
- Kimi K3 is now available in GitHub Copilot
- Show HN: The Channels SDK – Bring Any Agent to Any Channel (Slack, MS Teams)
- 🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp
- A llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs, 8B decode goes 2.39 → 8.20 tok/s
- Size the Orbs of Production!
- Show HN: A workflow for building community skill catalogs
- GitHub Code Quality no longer adds Copilot as a reviewer
- Managed Deep Agents is now in Public Beta
- llama.cpp PR reports up to 169% faster quantized-KV decode at 118K context on Intel Battlemage from one SYCL kernel switch
- Copilot usage metrics API adds agent app activity
- MCP allowlists in enterprise managed settings
- Copilot code review effort levels are generally available
- Qwen 3.6 27B flags/settings in llama.cpp
- GitHub Copilot weekly releases — August 3
- Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?
- I got tired of my 300GB model loads taking 5min on RPC. PR 26291 speeds it 300% to 1min30sec (4060ti+ddr4) + (4060ti+ddr5)
- Qwen 35B-A3B MoE vs 27B dense in local coding tests: ~4× faster, much smaller quality gap than I expected
- Qwen3.6 27B + 35B on vLLM, single R9700 (gfx1201)
- Claude Code in 9 lines python
- Building a zero-dependency C inference engine for BitNet (1.58-bit) - lessons from hitting 36 tok/s on a Xeon CPU
- Show HN: A terminal glued to the macOS dock
- Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size
- Extremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?
- ds4 flash 0731 UD-IQ2_M wrote a custom metal kernal for kimi k2 IQ1_0 in about 50 minutes
- Best Embedding + Reranking Model
- AMD llama.cpp: reducing MTP buffer overhead gave me 64K → 149K context for Qwen 27B
- DeepSeek v4 Flash 0731 locally on CPU
- Two flags took the official Ling-3.0-flash INT4 from 20.8 to 38.7 tok/s on one DGX Spark
- [NEW MODEL] SupraElegans-500K
- A Dial for You
- Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows
- 1M context with 17 GB model in 24 GB VRAM: "for the first time I was able to load a context of almost 1M tokens and extract 7 needles from various parts of the text"
- Engineers… Your Software Factory NEEDS Agent Sandboxes to SCALE (exe.dev)
- OpenAI Astra pause 🚨, Claude Code cross-session 🤖, how Cursor Router works 🔀
- Copilot on web expands conversation controls
- Best Local LLMs - August 2026
- Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
- DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks
- Show HN: Ante, a coding agent in a single binary that runs offline
- I compared GGUF quants of Qwen3.6 27B to NVFP4, AWQ, AutoRound, and FP8
- Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
- Meta Muse Glimmer 30B Local AI Review
- Global Plugins and Skills
- Show HN: Mcptoon – Token-efficient MCP CLI client
- Introducing Unsloth Desktop app
- You Don't need to use Cloud AI! Switchyard and Nemotron 3.5 Lightning
- How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation
- Automating and Scaling Behavioral Scientific Research on AI Agents
- FYI: Muse Glimmer Chat Template Got Updated Recently
- LangSmith BYOC is now generally available on AWS
- Google Maps and Google Search now work together in the Gemini API
- Introducing Delta
- Agent Plugins 1.0 in VS Code, Copilot CLI, and the Copilot app
- Meta's Muse Glimmer 30B now runs up to ~3.3x faster on Mac with mlx-dspark
- Why managed agents are the next big thing in agent building
- Show HN: Ballet – Workflow automation that writes integrations against any API
- Tool: Amp
- Beta: git-knife
- Beta: Docker Sandboxes
- Every Claude Code Skill I Use to Drive My Entire Development Process
- How do you plan to run Qwen3.8-2.4T-A95B locally?
- Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents
- Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
- The builder’s guide to GPT‑5.6
- Gemini 3.7 Flash is now available in GitHub Copilot
- Trained a 1.5B to write shell commands so I'd stop googling tar flags. Runs on a laptop CPU in ~1 sec.
- Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release
- Show HN: MCP Memory – Fast Agent Memory Using Google's OKF and SQLite FTS5
- Show HN: MCP-stama – An ultra-fast Rust MCP server with no dependencies
- LFM 2.5 2.6B is the best small model for tool use I have ever used.
- bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s
- fantastic: latest llama.cpp server webui can now run commands for tools into rootless sandboxed containers
- Senv: Sandboxed Python environments with the uv workflow
- Grok 4.6 is now available in GitHub Copilot
- GitHub Copilot weekly releases — August 10
- Local uncensored Opus 4.6 at home - Qwen3.8 27B heretic
- Show HN: Mole – Deep research agent for your terminal
- Show HN: Ember – Redshift safe color palettes
- Flownie – Open and Visual Data Workflow Platform with AI Agent Assistance
- Fable 5 refuses to touch Qwen deployments?
- Show HN: ThoughtDAG – An editable context graph for LLM conversations
- A nice local vision test
- SOTA Apple Silicon Inference (August 15, 2026)
- Show HN: Deltix – AI Driven Testing
- If you are at the lowest budget, which you can think of.Which hardware would you recommend to run? qwen 3.8 27b oWith like 50 tokens per second. I currently have a RTX 5070 Ti.
- Qwen3.8-27B Hybrid IQ4_XS quantization for 16GB gang
- Show HN: Laptop is the last place your secrets are still in plaintext
- The Architect: Interactive Visualization of Deep Learning Mathematics Directly in Microsoft Excel
- Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents
- Petition to add a rule for people to add their DAMN quant levels to their posts
- Ling 3.0 support merged into llama.cpp
- 100$ worth of gpu runs qwen 3.8 27b at 7.39 t/s
- Controlling Android with Gemini 3.7 Flash and 150 lines of Python
- After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)
- llama.cpp version v0.1.0 has been released
- Talk to Puck
- Same Cluster, 33 Points More Utilization: What Changed Was the Order
- llama.cpp adaptive MTP PR#27210
- Education Discount
- Qwen 3.8 27b saved me $650+ in API costs this evening
- Agentic Commerce at Scale: Your LangChain agents can transact securely
- Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
- Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB
- Introducing LangSmith Tuned Evaluators, starting with Perceived Error
- Asana cleared 5 years of engineering work in 2 weeks with Codex
- Claude Design artboard workflow added to Claude Code CLI
- Show HN: Runbook.v1 – governed workflow execution for MCP (fail-closed)
- Enterprise managed settings in GitHub Copilot for JetBrains
- Qwen3.8-27B on 2x 3090 + vLLM + DFlash2: 218 tok/s single request
- MCP in Orbs
- Replit expands access to software creation with GPT-5.6 Luna
- NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.
- Pass the Orb to the Left Hand Side
- DFlash2 speeds Qwen 3.8 27B up to 4 times
- Tool: TurboVec
- Beta: Saggar
- Beta: Needle
- DeepSeek Just Built the Next Generation of Coding Agents
- The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches
- TinySearch v0.6.1 - still a lightweight web research tool for local LLMs, now with bring-your-own-browser support
- How ChatGPT Work helps Stampli move ideas to market
- LangSmith Preview Builds: Test agent changes before production
- Unsloth Dynamic 3.0 GGUFs
- Show HN: Huzzah – a novel approach to coding with AI
- Qwen 3.8 27b - PI AGENT vs OPENCODE
- An Evidence-Grounded Multi-Agent System for High-Level Bio-Robot Design
- The Evaluation Context Protocol (ECP): A Portable Contract for AI Agent Evaluation
- Qwen3.8-27B at 262K context on a Strix Halo + RTX 3090 Ti: 9.5 -> 153 tok/s, and it beats a dual-3090 vLLM box on HumanEval
- Fastest NVFP4 quant of Qwen3.8 27B out there
- Explain Usage
- ChatGPT Apple Messages 💬, Anthropic’s meeting recorder 💼, Mistral Agentic Search 🔍
- DeepSeek Harness v0.1.1 released
- Filenames are the wrong index for Claude Code @ mentions
- Shared agentic work with GitHub Copilot in Microsoft Teams
- Getting the Same Results with Smaller "Cheaper" Dual Sparks AI as the More Expensive Clusters
- Qwen3.8-27B Q6 is a beast at agentic coding
- 16 GB VRAM purgatory discussion thread
- Show HN: OzBrain, a shared brain for knowledge between agents and your team
- Show HN: Omacosy – Omarchy-style tiling desktop for macOS, no SIP
- Show HN: Shoehorn – Quantize any model down to run on your machine
- This is why I run locally.
- Qwen 3.8 27b - PI AGENT vs OPENCODE - another smaple
- Llama.cpp version 0.2.0 is out!
- The New MCP Roadmap
- The Official Ruby SDK for MCP Reaches 1.0
- Show HN: Git workflow as an AI-agent skill
- I forked Ninfer 3090 and converted it to run on the CMP170HX - doubled my Qwen3.6-35B from llama.cpp
- GLM and I created a llama.cpp fork optimized for AMD GFX906 (Mi50, Mi60, Radeon VII, GCN HIP) - Machine Learning, LLMs, & AI
- I benchmark DFlash 2 (PR build) in llama.cpp on Qwen 3.8 27B against all speculative methods for 3 days. 2.26x on 100 real coding prompts, 4.68x with one n-gram drafter on top. Up to 8x on specific cases.
- Show HN: terminal-code – VS Code inside the terminal
- I fine tuned Gemma 4 12B for a 2.7x improvement on tool calling because I can't fit anything else comfortably into my 16 GBs of Vram
- I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens
- i finally switched from windows to linux and got a 30-50% boost in speed.
- Show HN: 26-node n8n workflow for scoring and routing B2B leads
- Qwen3.5-9B Triple-Loop
- Qwen 3.8 27B for actual local programming
- Show HN: Froging AI – image and video models in one workflow
- Friendly URLs for Sharing Orbs
- Qwen 3.8 27b helped me with something unique that Opus 4 couldn't - Firmware + Software preservation and emulation on an early 2000's ARM based POS system
- Live Artifacts: Authoring Dynamic Media via Live Layers Encapsulating Generative Specifications
- PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure
- I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB
- FreeToken Deepseek V4 Flash on a Single 3090 Local AI Testing
- Advancing price-performance for developers with GPT‑5.6 in Kiro
- JetBrains local AI (using Qwen3.6 27B)
- Please join r/LowEndLocalAI, a community for running local LLMs on low spec hardware
- Wire It, Run It, Deploy It: AI Workflows in Gradio
- I just tried DeepSeek Harness and it escaped from its workspace folder
- OptiMAS: Automatically Optimize Multi-Agent System
- Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments
- Show HN: Kern – container and resource runtime in a 1.5 MB binary, no daemon
- Do not blindly delete your older models, some are still precious
- Show HN: Screen memory without screenshots, just text to Markdown
- Every visual workflow tool becomes spaghetti. Here's what I built instead
- How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
- New: Llama.cpp adaptive speculation for faster inference
- New in LangSmith Engine: >2x better issue detection
- Introducing the Admin plugin for ChatGPT Work and Codex
- Show HN: I made a Raspberry with Qwen my local car AI
- GitHub Copilot app Customize tab is generally available
- Maiao: Gerrit-style code review workflow for GitHub, GitLab, Gitea, others
- Setup Without a Commit
- Whoever the fuck predicted we would have gpt 5.5 performance in coding on consumer hardware a couple months ago now, i applaud you
- The Future of SaaS Is Apps That Agents Can Use
- Show HN: Rudder – Red-Green TDD Workflow for Verifiably Comprehensive Specs
- August 2026: LangChain Newsletter
- Lemonade end-of-summer project update, now serving 15 engines!
- Beta: Huzzah
- Beta: Kern
- Beta: Apache Maka
- Beta: OpenViking
- Enterprise-managed settings now support autoUpdate for plugin marketplaces
- MCP-Driven Accessibility Tree Standardization for AI-Powered Screen Reader Agents
- So Long, TUI Sidebar
- llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp
- Projects with Multiple Repositories
- We’re the Team Behind Apodex 1.1 — Ask Us Anything!
- Request: unsloth Please re-quantize Qwen3.6 35 A3B and 27B using UD 3.0
- No, Engrams won't let you run 1T models locally. It does something even better.
- Show HN: My Claude quota ran out in 10 minutes, so I made a tool to find out why
- llama.cpp support for Qwen3.8-Flash-Next has been merged
- Copilot code review: Resolution reasons and expanded capabilities
- Actions retention will cover checks, workflow runs, and statuses
- Show HN: We built open OpenRouter that turns usage into a better model
- Previewing the Model Hardware Standard
- Amp on iOS & macOS
- Actions retention will cover checks, workflow runs, and statuses
- It's unbelievable! I used the mmap function in llama.cpp to fit Qwen3.8-Flash-Next IQ3_XSS into 16G+64G RAM, and the speed still reached 26t/s.
- ds4 branch with GLM 5.3 Flash support
- Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)
- No Mailmap Required
- Rime Finally Made Voice Agents Good Enough
- GitHub Copilot in Visual Studio — August update
- GitHub Copilot weekly releases — August 24
- A smarter way to run code migrations with less LLM context
- [Release] SOTA GGUFs for Qwen3.8-27B: GSQ-RCO at 2.5 to 3.0 bpw
- Is it worth running Qwen 3.8 Flash Next on 4x3090 vs 27B?
- Qwen 3.8 Flash Next ngram look up table offloaded to SSD and streamed in SGLang
- Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp)
- llama.cpp Open PRs list - CPU/RAM/Disk/Hybrid Related - Better for CPU-only & Hybrid inference
- Koboldcpp v1.120 released
- I fine-tuned a 0.8B local model for dictation cleanup. It matched a hosted frontier model on this narrow task
- Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and J-Wash Enhanced Fork!
- Don't Sleep on EXL3 Quants
- Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM
- Demo of local document extraction (52 pages) using Arctic Embed and Bonsai 8B on an Iphone 16 (KernelAI app)
- Qwen 3.8 Flash Next locally on simple mobile phone at 3.5 tok/s
- Here my pretty good qwen3.8 27B setup, hope it helps
- GOD: Govern, Observe, and Direct - A Real-Time Control Room for Agent Societies
- pipecat-ai/phonellm-alpha-1: GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost
- How I got Qwen 3.8 27b running at ~75t/s decode on 16GB RTX 5080
- Setup OpenClaw 2.0 with Gemini in Under 60 Seconds
- First time running local models
- GitHub Copilot in VS Code, August 2026 releases
- AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 · Pull Request #27402 · ggml-org/llama.cpp
- Show HN: Silent Fail, get an email when your n8n workflow stops running
- The Web-CLI: Verifiable Privacy for Tools, Models, and Inference Engines in the Browser
- Mac ← USB-C cable → Linux box is becoming a thing.
- ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++
- I pushed Qwen3.8-27B to 2.000 prefill per second and 132 decode per second on A RTX 3090.
- Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
- 11 Tiny Coding Agent Fixes With A Stupid Amount Of Payoff
- Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
- Intelligently Ordered Diffs
- Fable 5.1
- Selected GitHub Copilot models deprecated
- Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents -- A Source-Code Study of Eleven Systems
- CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
- ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s
- Confirmed bolting Q8 NGram into IQ4 Qwen no speed degradation
- Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
- Enterprise-managed settings support any default model
- Content exclusions generally available in Copilot app and CLI
- Perplexity open-sourced their Mac inference server for Qwen 3.6
- Agents That Pay | How Nevermined Empowers LangChain Agents to Buy and Sell Services
- Microsoft VibeVoice-ASR-Streaming Released
- Agent Flight Recorder: Tamper-Evident Audit Trails with On-Chain Anchoring for Long-Horizon Tool-Using Agents
- Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR
- KV cache might be a bigger problem for local models than parameter count
- model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support by YanissAmz · Pull Request #25444 · ggml-org/llama.cpp
- Agent-manager: The fastest workflow for every coding agent
- Can a 4B local model actually feel like an AI assistant?
- Give Your Coding Agents a Memory You Own
- Qwen-3.8-Next-Flash Ngram Hot-Swappable Knowledge Injector for llama.cpp
- Qwen3.6 35b Q2_XXS: Being GPU poor in 2026 is not so bad
- *NeoMME*: an efficient Multimodal-native and Multilingual Encoder
- We built an open-source, model-neutral agent harness and compared it with claude managed agents - for the same model, got same accuracy, upto 75% lower cost
- Muse Spark 1.3 ✨, Gemini 3.8 Flash ⚡, intelligence vs cost 📊
- MCP in LangChain: Stateless Protocol, Elicitation, and More!
- Playco cut manual fixes 50% prototyping games with GPT-6 Astra
- Upcoming deprecation of selected GitHub Copilot models
- GitHub Actions: Early September 2026 updates
- Desktop
- Model: add Tencent Hy 4 (hy_v4) preview architecture support by Little0o0 · Pull Request #28127 · ggml-org/llama.cpp
- GPT-6 Astra is generally available in GitHub Copilot
- GitHub Copilot weekly releases — August 31
- If you had ~15k would you build a home server today or wait
- OpenClaw Power, MacBook Simplicity: Five Days With Grok Bot
- NInfer vs llama.cpp vs vLLM: quality + speed comparison for Qwen3.8-27B NVFP4 on RTX 5090
- Otaku — an LLM frontend
Papers & ergonomics
- Central Tendency Bias in Human Selection of AI-Generated Design Variations
- Knowledge-Based Design Requirements for Generative Social Robots in Higher Education
- Advances in footwear and surface for prevention of occupational slips and falls - A scoping review
- Analysis of work system components in interprofessional communication to determine shock etiology
- The distracting role of stress: Impaired executive attention and delayed fatigue perception
- Vision-language models for occupational physical exposure assessment: Classification and temporal segmentation of manual material handling tasks
- Digital tools for identifying work tasks in various occupations: A scoping review
- Contributions of headset IPD fit, vection and sway to cybersickness during head mounted display based virtual reality
- Neurophysiological synchrony as an emergent performance marker of multi-human–robot team effectiveness
- A model for predicting pointing time in an eye-gaze input system using three basic phases of cursor movement trajectories
- Estimation of dynamic spinal loads during manual lifting using smartphone-based markerless motion capture
- Evaluating model-estimated shoulder muscle activity during overhead work with varied task demands and exoskeleton use
- Effects of sleep deprivation on cognitive-motor functions and adaptive skill learning among medical residents across 26h night shifts
- Population-specific ergonomic design and evaluation of a head-mounted display based on Chinese craniofacial measurements
- Import AI 456: RSI and economic growth; radical optionality for AI regulation; and a neural computer
- Effects of static postural loading on the performance of short-term/working memory tasks
- Gender and body height discriminate spinal movement patterns during lifting and lowering tasks
- Switching between touch and voice: factors influencing modality selection in multimodal systems
- Impact of anthropomorphism in AI assistants’ verbal feedback on task performance and emotional experience
- Evaluating biomechanical risks in manual material handling: an ergonomic intervention approach
- Effects of passive Arm-support exoskeleton on dynamic balance in different occupational tasks
- Learning to apply Design Thinking in participatory ergonomics: an exploratory study of OHS professionals in Denmark
- Recognising and explaining mental workload using low-interference method by fusing speech, ECG and eye tracking signals during simulated flight
- Size matters ̶ effect of screen setup on muscle activity and posture in computer work
- Leveraging socio-technical systems to tackle grand challenges: Reflections on human-robot teams, hybrid workplaces, med-tech, and digital transformation
- My coworker's 36 key Corne open-source keyboard setup
- When LLM Tutoring Responses Work: Evidence from Student Programming Conversations
- Same Stories, Different Journeys: From Social Comparison to Sensemaking in AI-Mediated Peer Career Exploration
- HandPad: A Bimanual Hand Interface for Fluid Window Interactions in VR
- Supporting Reflection in LLM-based Exploratory Search
- Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment
- TRAIL: A Platform for Configurable Human--AI Teaming Experiments
- AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized
- Role-specific experiences of inefficiency in the biomedical research and clinical laboratory workforce
- Conversational Tactile Data Interfaces: Co-Designing Accessible Data Experiences with Blind Users Using Refreshable Tactile Displays and Conversational AI
- When AI Blurs the Boundaries of Contribution: An Empirical Study of Authorship Calibration
- Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)
- Can AR Embedded Visualizations Foster Appropriate Reliance on AI in Spatial Decision-Making? A Comparative Study of AR X-Ray vs. 2D Minimap
- Initial static, pseudo-static, dynamic, and cognitive fit evaluation of three passive shoulder exoskeletons while performing simulated manufacturing tasks
- Human factors integration in complex systems: Awareness, challenges and strategies
- Smart glasses: environmental perception and perceived health effects
- System for activity-aware fatigue evaluation (SAFE) framework: Predictive fatigue modelling for occupational tasks
- It Matters How You Say It: Exploring Rhetorical Patterns for AI-Assisted Information Evaluation
- Informal Learning Emerges in Everyday Human-LLM Interaction
- Effectiveness of passive back-support exoskeletons during simulated commercial crab fishing tasks: Acute effects on muscle activity, joint kinematics, and subjective measures
- Effects of uncertain system information on users' performance, confidence, and trust in aided linguistic judgement
- Toward context-aware and personalized sit-stand desk interventions: Insights from a field observational study
- Textual cues, cognitive load, and social fatigue: Unveiling the reasons behind user discontinuance in conversational AI
- Knowledge Gaps and Explanation Design in AI Advice: A Cognitive Fit Perspective on User Decision Confidence
- The creative tax of videoconferencing: How ideological diversity mitigates creativity losses in virtual teams
- Trust erodes, fatigue builds: How prompt uncertainty traps users in recrafting loops
- The Digital Mindfulness Scale: Development and longitudinal validation in the workplace
- A Comprehensive Evaluation of Job Rotation: Biomechanical Risk, Body Discomfort, and Psychosocial Demands
- Comparing Paper- and AR-Based Assembly Manuals in Task Performance, dlPFC Hemodynamic Responses, and Perceived Workload
- The Influence of Exposure and Error Type on Estimates of Automation Reliability
- Mind-Wandering or Task-Unrelated Thought Reports May Be a Response to Performance Not a Cause of Performance: Using Forced Errors to Impact Thought Content Reports
- Vigilance Research Beyond the Laboratory: Methodological Considerations and Practical Insights
- Efficiency Pitfalls of Explainable AI in Clinical Diagnostic and Treatment Human-AI Workflows
- Human–Artificial Intelligence Collaborative Decision-Making in Emergencies: Relative Advantage Theory
- Modeling Human Expertise in a Sanding Task
- Compatibility Effects With Simple Lever Tools: A Replication and Extension Beyond Simple Button Responses
- System-Wide Trust (SWT) Versus Component-Specific Trust (CST) in Multi-Agent Human–Agent Teams: Individual Variability in Trust Bias
- Effects of Task Priority and Difficulty in Multitasking Across Screens
- Ethical Decision-Making in the Workplace: Which Factors Influence Decision-Makers’ Willingness to Revise Their Decisions due to AI-Generated Suggestions?
- Decision-Supported Visual Search: Effects of Direct and Indirect Cues on Performance and Attentional Mechanisms
- A Structural Model of Attentional Effort Dynamics: Evidence From a Naturalistic Discrimination Task
- Affective Interruptions Impair Task Performance Less Than Neutral Ones Across Age and Task Demands
- Comprehensive Evaluation of Explanation Types in a Spaceflight-Relevant Human–Autonomy Teaming Task
- What’s in a Name? Implications of AI Roles and Mind Perception for Human-AI Teams
- Can We Learn From Trust in Simulation to Gain Trust in AI?
- Systemic Trust in Artificial Intelligence
- Does Expertise Matter? A Study of Chile’s Ergonomic Risk Assessment Tool
- Effectiveness of Eye-Tracking Metrics in Human-Centric Design of Human-Machine Interface: Cases on Process Control Operations
- Three Participatory Methods to Engage Employees in Workplace Research and Design
- An Exploratory Study of the Causes of Discomfort From Using Braille Displays
- Spatial Discontiguity in Three-Dimensional Augmented Reality Spaces
- Satellite Ground Stations: The Need for Additional Research and Standardization
- An Empirical Study of Tablet Ergonomics: The Interplay of Temperature, Orientation, and Use Behaviors
- Contours of Comfort: Mapping Pressure Landscapes Across Anthropometric Seating Interfaces
- Workplace Ergonomics in Bangladesh: A Scoping Review of Anthropometric Mismatches and Musculoskeletal Disorders
- Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists
- TargetFinder: Detecting Widgets from Pixels on Desktop Interfaces
- Magnus Evo XL review: Height-adjustable gaming desk put to the test
- Mitigating psychological discomfort during automation error: A mixed methods research
- CRAFT: Exploring Wearable Creative AI on Smart Glasses for Fiction Writing in Real-World Contexts
- Transparent by Design, Usable in Practice? A Formative Usability Study of a Conversational Product Advisor
- Delegating (or Not) to Machines: How Role Expectations Shape Leaders’ Willingness to Use AI for Communication
- Bespoke Visual Assistance: What and How do Blind and Low-Vision People Create with Agentic Programming?
- A Method for Constructing a Dynamic Seat Adjustment Model to Mitigate Prolonged-Sitting Discomfort
- Touching or Chatting: The Utility of LLMs and Tactile Charts for Learning about Complex Chart Types by BLV Individuals
- Relationships Between Trust, Compliance, and Performance for Novice Programmers Using AI Code Generation
- Verification Without Distrust: Reframing User-Side Oversight as Routine Epistemic Governance in Everyday Human-Chatbot Interaction
- Using Eye Tracking to Identify Comfort Attributes in Office Chairs
- How Wrangling Tools Shape Wrangling: A Technical Dimensions Analysis
- FleetScape: A Mixed Reality Sandtable for Spatial Supervision and Control of Scalable Drone Fleets
- Evaluating the Vergence-Accommodation Conflict in Gaze-Based 3D Target Selection
- Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
- $\Sigma$-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems
- More externalization, but less inference? Exploring changes in young learners’ critical thinking during conversational AI bot-supported multimodal writing practice
- An experimental investigation on the relationship between seat pan and seat back angles to eliminate the shear force on the seat pan
- Cognitive Readiness for Human-AI Collaboration
- Visualizing Placement Proposals for Window Arrangement in Mixed Reality: A Comparative User Study
- Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits
- Efficient Optimal Mouse Sensor Position Estimation using Simulated Cursor Trajectories
- Revisiting Channel Effectiveness: A Multi-Dimensional Evaluation with Primitive Visual Stimuli
- Mixed Uncertainty in One View: Co-Visualizing Statistical Variability and Qualitative Confidence
- Toward Resilient Human-AI Collaboration: A Lifecycle Taxonomy of Sociotechnical Risks and Cascading Failures
- CaRing: Preventing Carpal Tunnel Syndrome based on Daily Activities from Always-Available Input Device
- Topic Matters: How Linguistic Properties can Shape Reading Behaviour in Selective Exposure Studies
- Obesity-related differences in shoulder and lumbar loading during manual material handling
- Effect of non-neutral seated postures on whole-body vibration measurements
- Not Always Top-Left: Untangling the Signals that Guide Dashboard Reading Order
- Ergonomic office chair with lumbar support and footrest: Welax S9 Pro hands-on review
- Large Language Models Explain Experts Better Than Experts Themselves
- Concurrent validity of markerless motion capture for occupational lifting: Kinematics and ergonomic risk assessment
- Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents
- TRIBE: Predicting Team Performance via Communication Behavior Ensembles
- How Do Senior Qualitative Researchers Perceive the Risks and Opportunities of Using Large Language Models in Qualitative Analysis? An Exploratory Study
- Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
- From “cost of asking” to “fit of asking”: How seeking help with AI shapes employees’ indebtedness and autonomy in different workplace helping contexts
- Ontology-Grounded World Models for Failure Diagnosis and Closed-Loop Repair in Physical AI Systems
- How Cognitive Load Affects Dynamic Trust Calibration in Human–AI Collaboration: Evidence for Selective Pathway Effects
- YouthShape.US: A New Anthropometric Resource for U.S. Children and Youth
- What Cognitive Accessibility Reveals About Data Visualization
- Spatial alignment of audio–haptic directional feedback shapes non-visual hand guidance
- Chat First, Worry Later: Understanding Individuals' Privacy Perceptions Using ChatGPT in a Work Context
- What is the current state of evidence on passive exoskeletons designed to support the lumbar region? A systematic review and meta-analysis of physiological, biomechanical, and user-related outcomes
- Assessment of pressure discomfort for over-ear wearables and the relationships with objective metrics
- ‘I’d feel like management understands (no pun intended) how we feel’: evaluating a hypothetical policy promoting sitting in standing-biased jobs
- Effects of prismatic loupes on surgeons’ intraoperative physical workload and musculoskeletal discomfort in operating room
- Validating force-estimating insoles for calculating centre of pressure and vertical ground reaction forces during occupational tasks
- Mastering a robot workforce: review of single human multiple robots systems and their impact on occupational safety and health and system performance
- Automatic estimation of Hand Activity Level from upper-limb trajectories: a probabilistic regression framework
- State of science: New frontiers in inclusive design and digital health intervention
- Physiological measurement of situation awareness: a study of the validity of EEG and fNIRS during performance and automation monitoring in a complex task
- Effects of dynamic lighting on neurobehavioral performance under different mental states in the working area of a space station
- Aura: Dynamic Intra-Turn Emotion-Aware Adaptation of Large Language Model Responses
- Using physiological measures to assess and diagnose team performance: An interrogative approach
- Investigating Human Factors and Ergonomics research: a 4S framework
- Applying cognitive and perceptual science to typeface choices
- A scoping review on emerging technologies and automation of musculoskeletal ergonomic assessments
- Strategic music listening and subjective perceptions: effects on attention and workload across task complexity levels
- Enhancing mental workload recognition: a comparison of complexity-based eye movement metrics and conventional features
- An ergonomic intervention to minimise physical and physiological stresses in the office standing workstation
- REBA integrated with organisational analysis to assess the risk of biomechanical overload in physiotherapists
- The effects of prolonged standing in occupational footwear on perceived discomfort, standing balance, and gait biomechanics in young adults
- RegulAR: Graph-Grounded Error Recognition and Assistance for Procedural Tasks in AR
- Dynamic Tree Colors: Adaptive Discriminable Hierarchies with Minimum Instability
- Graphionale: How Graph Visualizations of LLM Rationales Affect Human Decision Making
- Too Much of the Same: From Algorithmic to Human Bias in Learning to Defer
- User Preferences for UI Anchoring in MR: Effects of Task Mobility and Interface Properties
- Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI
- Toward Postural State Classification in Immersive VR with Multimodal Data and Explainability Analysis
- Collaboratively Eliciting Gestures for Geospatial Data Exploration on an MSE with Tangibles and Styluses
- FocusBuddy: Encouraging Healthy Desk-Work Habits by Caring for a Virtual Pet on a Water Bottle
- ErgoAssist: Cognition-Aware Posture Feedback in Wearable Ergonomic Systems
- RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces
Practical tips
- Giving your AI a Job Interview
- Layout Buffet - Mousing with a keyboard
- Layout Buffet - Layers
- Practice typing with your favorite book excerpts
- Layout Buffet - Home-row Mods
- Layout Buffet - Sticker Mods
- Automate Excel with Python: From manual grind to one-click workflow
- Brave's latest browser release offers Containers for better and easier workflow
- Made a free macOS menu bar app that fixes typing in the wrong keyboard layout
- Mouseless – keyboard-driven control of macOS/Linux/Windows
- Neverclick: Desktop application for performing mouse actions with your keyboard
- Getting started with ChatGPT
- Gemini 3 Prompting: Best Practices for General Usage
- Hidden Gems: Part 3
- Show HN: I was tired of opening 2 tabs for every HN link, so I made a userscript
- PSA for DeepSeek-V4-Flash-0731 users — don't blow out your prompt cache with system role messages mid-conversation
- You really should not quantize KV Cache for DeepSeek V4 Flash
- How to Decide When an AI Tool Is Worth Keeping
- My Terminal Workflow for Note-Taking, Data Engineering and Writing (Linux/macOS)
- One AI Output Is an Example, Not an Evaluation
- Weirdly, no one talks about Temperature setting for the Qwen3.8 27b
- GPU Poor - Don't overlook Laguna XS 2.1
- Don't sleep on Vision support for coding!
- Everyone is t/s maxing.. 3.8.. but after a week of using it for work I'm tempted to switch back to 3.6
- My RULE of Thumb of choosing a models
Raw data: signal-graph.json
Signal. An evidence stream on how AI is changing the workday — sources read, scored, and filed.
read the stream →- Score 4.20curatedAgentic workflow patterns
The Illusion of Independent Quorums: Epistemic Fault Domains and Correlated Cognitive Failures in Agentic Quorums
The paper introduces Epistemic Fault Domains (EFDs) and a structural cut metric κ_E to formalize the failure mode where multi-agent quorums share upstream inputs, telemetry, or tool backends, collapsing multiple votes onto a single corrupted cause. It proves quorum size does not guarantee epistemic redundancy and presents the DAQC controller plus a 120-task benchmark.
arXiv cs.MA · T1
- Score 3.80curatedAgentic workflow patterns
The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems
Proposes the Civilization Framework, in which the addressable unit is a 'civilization' (one human sovereign, a persistent ledger, interchangeable agents) rather than individual agents, with an Embassy Protocol for asynchronous inter-agent message delivery. A preregistered 1,908-trial experiment reports a temporal-weight effect: incorrect upstream claims arriving first captured 54.2% of receiver answers. Results are flagged exploratory.
arXiv cs.MA · T1
- Score 4.00curatedAgentic workflow patterns
Where Reliability Lives: Experimental Localisation of Behavioural Properties in an Agent System
ArXiv paper experimentally localizing reliability properties in an agent system built around an append-only ledger adjudicator. Interventions on institutional epistemic mechanisms and on cognition (ablation, mid-task reset, frontier-LLM substitution, false testimony) left five core properties intact: singular accepted reality, typed refusals, durable duties, no double-acceptance, and no false completions across 2,581 substituted-panel claims.
arXiv cs.MA · T1
The ledger’s credibility
How the Signal ledger earns trust.
Four rules keep every reading auditable — from where a claim comes from to how the map is allowed to draw it.
Tiered, reviewed sources
Every source sits on a reviewed list. Its tier weights the final score by provenance — an auditable code weight, not a trust badge.
Computed, not conjured
Five dimensions are scored, then combined by formula — mean × source weight. The final always says what it is, never a model’s opinion.
Source and reading kept apart
What a source said and what the pipeline read never merge into one voice. Pipeline notes always carry their disclaimer.
Machine guesses stay labelled
On the map a dashed link is an embedding’s suggestion; a solid link is an editor’s citation. The two never blur.
Tile values are illustrative examples — not a live entry’s reading.











