最近更新
更新:2026-09-12 12:06:22 · 公开链接
Atlassian Rovo A2A:真实 Agent Card、企业授权门与流式任务边界ai/sources
wiki/ai/sources/atlassian-rovo-a2a-live-card-and-enterprise-gates.md
BenchShield:从隔离 verifier 到奖励全生命周期完整性ai/sources
wiki/ai/sources/benchshield-reward-lifecycle-integrity.md
企业 MCP 鉴权网关:persona × credential 与服务账号授权边界ai/sources
wiki/ai/sources/enterprise-mcp-persona-credential-boundaries.md
llm-wiki 自我优化 2026-09-12:实现不覆盖标准,高分不覆盖弃权ai/sources
wiki/ai/sources/llm-wiki-optimization-2026-09-12.md
llm-wiki 自我优化 2026-09-11:可回读、可相信、获授权不是一回事ai/sources
wiki/ai/sources/llm-wiki-optimization-2026-09-11.md
全部页面
按知识分区浏览
ai 454
Loop Engineering vs Harness Engineeringai/comparisons
wiki/ai/comparisons/Loop-Engineering-vs-Harness-Engineering.md
Agent Evaluation & Harness Engineeringai/concepts
wiki/ai/concepts/Agent-Evaluation-and-Harness-Engineering.md
External Agent Skills Design Patternsai/concepts
wiki/ai/concepts/External-Agent-Skills-Design-Patterns.md
LLM Training From Scratch Education Stackai/concepts
wiki/ai/concepts/LLM-Training-From-Scratch-Education-Stack.md
Mediated Communication Can Steer Collective Opinion Generativeai/entities
wiki/ai/entities/Mediated_Communication_Can_Steer_Collective_Opinion_Generative.md
10x-bench-kit:企业内部 coding-agent benchmark 模板ai/sources
wiki/ai/sources/tenx-bench-kit-internal-agent-benchmark.md
A Recipe for Training Neural Networksai/sources
wiki/ai/sources/A_Recipe_for_Training_Neural_Networks.md
A2A Samples:Agent Card / Discovery / 多框架 agent interoperabilityai/sources
wiki/ai/sources/a2a-samples-agent-card-discovery-interoperability.md
A2A TCK:把协议兼容拆成实际覆盖、任务生命周期和安全边界ai/sources
wiki/ai/sources/a2a-tck-conformance-coverage-and-security-boundaries.md
A2Apex — Agent Card certification and discoverable A2A trust directoryai/sources
wiki/ai/sources/a2apex-agent-card-certification-directory.md
AAABench — long-horizon game-world coding-agent benchmarkai/sources
wiki/ai/sources/aaabench-long-horizon-game-world-benchmark.md
ACQUIRE — QA-Driven Repository Knowledge Acquisitionai/sources
wiki/ai/sources/acquire-qa-driven-repository-knowledge.md
ADeptS-Bench:跨设备 Computer Use Agent 可信度评测ai/sources
wiki/ai/sources/adepts-bench-trustworthy-computer-use-agents.md
AI Dev Team — supervised multi-agent delivery harnessai/sources
wiki/ai/sources/ai-dev-team-supervised-delivery-harness.md
AI 生成代码 Review 调研 2026-07-01ai/sources
wiki/ai/sources/ai-code-review-agent-coding-research-2026-07-01.md
AI-Mediated Communication Can Steer Collective Opiai/sources
wiki/ai/sources/AI-Mediated_Communication_Can_Steer_Collective_Opi.md
AISI SandboxEscapeBench:容器逃逸能力评测ai/sources
wiki/ai/sources/aisi-sandboxescape-bench-container-breakout.md
AIShield — Agent tool/content security scannerai/sources
wiki/ai/sources/aishield-agent-tool-security-scanner.md
AKM Eval — agentic knowledge management maturity indexai/sources
wiki/ai/sources/akm-eval-agentic-knowledge-management.md
AOI / DynaCU-Bench — dynamic computer-use observation interfaceai/sources
wiki/ai/sources/aoi-dynacu-bench-dynamic-computer-use.md
ASIL:用 structured state 与 semantic actions 替代 screenshot-and-clickai/sources
wiki/ai/sources/asil-structured-state-semantic-actions.md
Action-Graded Severity Scale for Tool-Using AI Agentsai/sources
wiki/ai/sources/action-graded-severity-scale-tool-agents.md
Active-SWE — proactive bug-fixing benchmark without issue reportsai/sources
wiki/ai/sources/active-swe-proactive-bug-fixing-benchmark.md
Aegis — architecture-aware coding-agent method packai/sources
wiki/ai/sources/aegis-architecture-aware-method-pack.md
AegisFlow Local-First Policy Gatewayai/sources
wiki/ai/sources/aegisflow-local-first-policy-gateway.md
Aftermath — execution-backed verification receipts for coding agentsai/sources
wiki/ai/sources/aftermath-execution-backed-verification-receipts.md
Agent Belt — black-box CLI agent benchmarkai/sources
wiki/ai/sources/agent-belt-black-box-cli-agent-benchmark.md
Agent Config:机器校验的跨宿主 skill/rule/command 库ai/sources
wiki/ai/sources/agent-config-machine-checked-skill-library.md
Agent Graph — fact-routed work contracts for Agent Skillsai/sources
wiki/ai/sources/agent-graph-fact-routed-skill-workflows.md
Agent Harness Evolution Shapes Coding Agent Qualityai/sources
wiki/ai/sources/agent-harness-evolution-shapes-coding-agent-quality.md
Agent Skills Hub and skills.sh High-Star Skillsai/sources
wiki/ai/sources/Agent-Skills-Hub-and-Skills-sh-High-Star-Skills.md
AgentBattler Bench — sealed harness benchmark for coding agentsai/sources
wiki/ai/sources/agentbattler-bench-sealed-harness-benchmark.md
AgentDoctor — coding-agent configuration auditai/sources
wiki/ai/sources/agentdoctor-coding-agent-config-audit.md
AgentFlow:面向 LLM Agent 的 flow-centric policy languageai/sources
wiki/ai/sources/agentflow-flow-centric-agent-security-policy.md
AgentFootprint:可查询的 agent 上下文因果轨迹ai/sources
wiki/ai/sources/agentfootprint-explainable-agent-traces.md
AgentInterdict — Runtime Authority Boundaryai/sources
wiki/ai/sources/agentinterdict-runtime-authority-boundary.md
AgentKernelArena:GPU Kernel Agent 的 A/B 与 RL-ready 评测环境ai/sources
wiki/ai/sources/agentkernelarena-ab-rl-gpu-kernel-agents.md
AgentLint — runtime guardrails for AI coding agentsai/sources
wiki/ai/sources/agentlint-runtime-guardrails.md
AgentLock — provenance-based pre-action authorizationai/sources
wiki/ai/sources/agentlock-provenance-action-gate.md
AgentMesh Runtime Gateway:task envelope 到 sandbox/telemetry receipt 的最小控制面ai/sources
wiki/ai/sources/agentmesh-runtime-gateway-task-envelope-sandboxclaim.md
AgentOps coding-agent verification membraneai/sources
wiki/ai/sources/agentops-coding-agent-verification.md
AgentOps — bounded coding-agent operating loop and skills bundleai/sources
wiki/ai/sources/agentops-bounded-operating-loop.md
AgentRelBench — stochastic damage benchmark for action-taking agentsai/sources
wiki/ai/sources/agentrelbench-stochastic-agent-damage-benchmark.md
AgentRoom:CRDT shared workspace for concurrent coding agentsai/sources
wiki/ai/sources/agentroom-crdt-shared-workspace-coding-agents.md
AgentTether — graph-guided runtime repair for LLM agentsai/sources
wiki/ai/sources/agenttether-graph-guided-runtime-repair.md
AgentTrust:工具调用的实时安全拦截与 TrustReportai/sources
wiki/ai/sources/agenttrust-runtime-safety-interception.md
Agentfootprint — contextual-error evidence graph for agentsai/sources
wiki/ai/sources/agentfootprint-contextual-error-evidence-graph.md
Agentic coding and persistent returns to expertiseai/sources
wiki/ai/sources/claude-code-expertise.md
Aperture — agent daily report curation skillai/sources
wiki/ai/sources/aperture-agent-report-radar-skill.md
Approving — human-gated multi-agent workflowsai/sources
wiki/ai/sources/approving-human-gated-multi-agent-workflows.md
Archestra — enterprise AI/MCP control planeai/sources
wiki/ai/sources/archestra-enterprise-ai-control-plane.md
AssetOpsBench — industrial agent benchmark over MCPai/sources
wiki/ai/sources/assetopsbench-industrial-agent-benchmark.md
Atlassian Rovo A2A:真实 Agent Card、企业授权门与流式任务边界ai/sources
wiki/ai/sources/atlassian-rovo-a2a-live-card-and-enterprise-gates.md
AutoDev Studio — repository knowledge base as multi-agent SDLC harnessai/sources
wiki/ai/sources/autodev-studio-repo-kb-sdlc-harness.md
AutoResearch 自动研究框架,通过 metric 搜索和 eval 三个关键动作,把赌变成ai/sources
wiki/ai/sources/AutoResearch_自动研究框架通过_metric_搜索和_eval_三个关键动作把赌变成.md
Autonomous Coding Agents 的 Security Debtai/sources
wiki/ai/sources/security-debt-autonomous-coding-agents.md
AxisAgentic — runtime and trajectory collection framework for long-horizon agentsai/sources
wiki/ai/sources/axisagentic-runtime-trajectory-framework.md
BOSS Console — agent operator console and governed workspaceai/sources
wiki/ai/sources/bossconsole-agent-operator-console.md
Baseerat:不可见屏幕时的 Computer-Use Agent 监督缺口ai/sources
wiki/ai/sources/baseerat-oversight-gap-computer-use.md
BenchShield:从隔离 verifier 到奖励全生命周期完整性ai/sources
wiki/ai/sources/benchshield-reward-lifecycle-integrity.md
Better Harness — evidence-bounded coding workflow reviewai/sources
wiki/ai/sources/better-harness-evidence-bounded-workflow-review.md
Beyond Pass@k:Agentic Code Generation 的可靠性与安全评测ai/sources
wiki/ai/sources/beyond-pass-k-agentic-code-reliability-security.md
Boundary-Bench — sandbox policy benchmark for coding agentsai/sources
wiki/ai/sources/boundary-bench-sandbox-policy-benchmark.md
BreachForge — agentic exploit containment harnessai/sources
wiki/ai/sources/breachforge-agentic-exploit-containment-harness.md
Bromure Agentic Coding:VM 边界上的 coding-agent 沙箱ai/sources
wiki/ai/sources/bromure-agentic-coding-sandbox.md
CUArena:把真实应用重建成可重置、可观察、可评分的 computer-use 环境ai/sources
wiki/ai/sources/cuarena-instrumented-computer-use-environments.md
Caliper — Reliability testing for agent skillsai/sources
wiki/ai/sources/caliper-skill-reliability-testing.md
Chancery — agent identity and writ-based controlai/sources
wiki/ai/sources/chancery-agent-identity-writ-control.md
Claimproof — evidence-gated final claims for coding agentsai/sources
wiki/ai/sources/claimproof-evidence-gated-agent-claims.md
Claw-Eval — trustworthy evaluation of autonomous agentsai/sources
wiki/ai/sources/claw-eval-trustworthy-agent-evaluation.md
CodeRail — convergent coding governance for agentic projectsai/sources
wiki/ai/sources/coderail-convergent-coding.md
Context Engineering Kit — Agent Skills Marketplaceai/sources
wiki/ai/sources/context-engineering-kit-agent-skills.md
ContextIQ — AST code graph for token-efficient agent contextai/sources
wiki/ai/sources/contextiq-ast-code-graph-agent-context.md
ContextNest / ContextNext — Verifiable Context Governanceai/sources
wiki/ai/sources/contextnest-verifiable-context-governance.md
ContextPilot:面向长程 Agent 的主动上下文管理ai/sources
wiki/ai/sources/contextpilot-proactive-context-management.md
ControlKeel — governed coding-agent control planeai/sources
wiki/ai/sources/controlkeel-governed-agent-control-plane.md
CooperBench — Cooperative Coding Agents Benchmarkai/sources
wiki/ai/sources/cooperbench-cooperative-coding-agents.md
Cua:computer-use drivers、sandboxes 与 Cua-Benchai/sources
wiki/ai/sources/cua-computer-use-drivers-sandboxes-benchmarks.md
CyVisGuard — MCP security control plane for AI agentsai/sources
wiki/ai/sources/cyvisguard-mcp-security-control-plane.md
Cyclops — deterministic MCP toxic-flow proxyai/sources
wiki/ai/sources/cyclops-deterministic-mcp-toxic-flow-proxy.md
DM-Code-Agent:本地优先、可审计、可分叉重跑的 coding agentai/sources
wiki/ai/sources/dm-code-agent-auditable-local-code-agent.md
Data Eng Bench — data-engineering benchmark for coding agentsai/sources
wiki/ai/sources/data-eng-bench-domain-coding-agent-benchmark.md
Deep Neural Nets 33 Years Ago and 33 Years From Nowai/sources
wiki/ai/sources/Deep_Neural_Nets_33_Years_Ago_and_33_Years_From_Now.md
DeepSWE — Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasksai/sources
wiki/ai/sources/deepswe-long-horizon-coding-agent-benchmark.md
Designer Skills Pack — agentic design skills and commandsai/sources
wiki/ai/sources/designer-skills-agentic-design-skill-pack.md
Deterministic Gates for Tool-Using Agent Policy Enforcementai/sources
wiki/ai/sources/deterministic-gates-tool-agent-policy.md
DexJoCo: A Benchmark and Toolkit for Task-Orientedai/sources
wiki/ai/sources/DexJoCo_A_Benchmark_and_Toolkit_for_Task-Oriented.md
Dispatch-Level Instrumentation:Agentic 抽取评测的工具轨迹证据ai/sources
wiki/ai/sources/dispatch-level-instrumentation-agentic-datasheet-extraction.md
Distributed Attacks in Persistent-State AI Controlai/sources
wiki/ai/sources/distributed-attacks-persistent-state-ai-control.md
Do Agent Optimizers Compound? — Terminal-Bench 2.0 持续学习评测ai/sources
wiki/ai/sources/agent-optimizers-compound-terminal-bench.md
Docker socket 沙箱旁路:真正的边界包括被委托的宿主 daemonai/sources
wiki/ai/sources/pillar-docker-socket-sandbox-trust-handoff.md
EVOMAL:自演化 coding agent 的 skill-library self-poisoningai/sources
wiki/ai/sources/evomal-self-poisoning-agent-skill-libraries.md
Enola — deterministic codebase architecture graphai/sources
wiki/ai/sources/enola-deterministic-codebase-architecture-graph.md
EvalGlass — local-first agentic app evaluationai/sources
wiki/ai/sources/evalglass-local-first-agentic-app-evaluation.md
Everdict — harness-agnostic agent evaluation runtimeai/sources
wiki/ai/sources/everdict-harness-agnostic-agent-eval-runtime.md
EvoSOP — Iterative Tool Optimization for Self-Evolving LLM Agentsai/sources
wiki/ai/sources/evosop-iterative-tool-optimization.md
FastContext — read-only delegated repository exploration agentai/sources
wiki/ai/sources/fastcontext-read-only-repository-exploration.md
Flow-Next — repo-native agentic engineering workflowai/sources
wiki/ai/sources/flow-next-agentic-engineering-workflow.md
ForgeOS — skill intelligence and trust control planeai/sources
wiki/ai/sources/forgeos-skill-intelligence-control-plane.md
From Prompts to Contracts — Auditable Enterprise Harness Engineeringai/sources
wiki/ai/sources/from-prompts-to-contracts-harness-engineering.md
GSD Pi — local-first agentic project workflowai/sources
wiki/ai/sources/gsd-pi-local-first-agentic-project-workflow.md
Golf Scanner:MCP server inventory 与 security auditai/sources
wiki/ai/sources/golf-scanner-mcp-server-security-audit.md
Grafana MCP:会话不是身份,凭据隔离不是出站隔离ai/sources
wiki/ai/sources/grafana-mcp-session-identity-and-egress-boundaries.md
Graph Engineering — knowledge graph and task graph topology for agentsai/sources
wiki/ai/sources/graph-engineering-agent-topology.md
Greplica:持久仓库记忆的 coding-agent 规划评测ai/sources
wiki/ai/sources/greplica-persistent-coding-agent-memory.md
HOL Guard — AI agent antivirus and runtime protectionai/sources
wiki/ai/sources/hol-guard-ai-agent-antivirus.md
Halo Record — tamper-evident runtime records for AI agentsai/sources
wiki/ai/sources/halo-record-runtime-records.md
Harness Arena:盲评 Agent Harness 而非模型ai/sources
wiki/ai/sources/harness-arena-blind-agent-harness-benchmark.md
Harness Engineering Anthology — repository as agent context bundleai/sources
wiki/ai/sources/harness-engineering-anthology.md
Harness Handbook:让演化中的 Agent Harness 可读、可导航、可编辑ai/sources
wiki/ai/sources/harness-handbook-evolving-agent-harnesses.md
Harness Score — deterministic maturity scanner for coding-agent harnessesai/sources
wiki/ai/sources/harness-score-maturity-scanner.md
HarnessRouter 与 Unified Harness Protocol:自托管 agent harness 网关ai/sources
wiki/ai/sources/harnessrouter-unified-harness-protocol.md
HealthClaw Guardrails — clinical data MCP guardrail layerai/sources
wiki/ai/sources/healthclaw-guardrails-clinical-agent-data-boundary.md
Hermes Field Kit — field-tested skill admission and validation contractai/sources
wiki/ai/sources/hermes-field-kit-skill-admission-contract.md
Hermes MemConflict memory-provider benchmarkai/sources
wiki/ai/sources/hermes-memconflict-memory-provider-benchmark.md
Higress v2.2.4:MCP、Gateway API 与推理扩展ai/sources
wiki/ai/sources/higress-mcp-gateway-api-inference-extension.md
IBM ContextForge:MCP / A2A / REST 联邦控制面ai/sources
wiki/ai/sources/ibm-contextforge-mcp-federation-control-plane.md
IssueBenchKit:把真实 GitHub issue 变成私有 coding-agent benchmarkai/sources
wiki/ai/sources/issuebenchkit-private-repo-coding-agent-benchmark.md
Karpathy 是 AI 领域知名研究者,前 Tesla AI 总监,前 OpenAI 联合创始人ai/sources
wiki/ai/sources/Karpathy_是_AI_领域知名研究者前_Tesla_AI_总监前_OpenAI_联合创始人.md
Keidai:MCP Gateway、Agent Identity 与 Operator Control Planeai/sources
wiki/ai/sources/keidai-mcp-gateway-identity-control-plane.md
Kitbash — cross-host agent skill format and compilerai/sources
wiki/ai/sources/kitbash-cross-host-agent-skill-format.md
LLM Wiki 是 Karpathy 提出的概念,给个人 Wiki 加 LLM 加持。与 Wikiai/sources
wiki/ai/sources/LLM_Wiki_是_Karpathy_提出的概念给个人_Wiki_加_LLM_加持与_Wiki.md
LLM-as-a-Verifier:细粒度 agent verifier 与 test-time scalingai/sources
wiki/ai/sources/llm-as-a-verifier-fine-grained-agent-verifier.md
Laborant Skills — evaluation design for AI systemsai/sources
wiki/ai/sources/laborant-evaluation-design-skills.md
Lilian Weng — Harness Engineering for Self-Improvementai/sources
wiki/ai/sources/lilian-weng-harness-engineering-self-improvement.md
Lilian Weng — Scaling Laws, Carefullyai/sources
wiki/ai/sources/lilian-weng-scaling-laws-carefully.md
LongHorizon-Harness — long-horizon computer-use state harnessai/sources
wiki/ai/sources/longhorizon-harness-long-horizon-computer-use.md
LongHorizon-Harness:面向真实电脑环境的长程 Loop Engineeringai/sources
wiki/ai/sources/longhorizon-harness-computer-use-loop.md
LongPuzzleBench:长程视觉 puzzle GUI agent benchmarkai/sources
wiki/ai/sources/longpuzzlebench-long-horizon-gui-agent-benchmark.md
Loop Engineering 实战:从日志扫描到预发部署的全自主闭环ai/sources
wiki/ai/sources/loop-engineering-autonomous-log-to-staging.md
Lovable — Scaling Agentic Coding with Token Spendai/sources
wiki/ai/sources/lovable-scaling-agentic-coding.md
MCP Guardrail — SQL authorizer boundaryai/sources
wiki/ai/sources/mcp-guardrail-sql-authorizer-boundary.md
Microsoft Agent Skills — context-driven development skill catalogai/sources
wiki/ai/sources/microsoft-agent-skills-context-driven-development.md
Microsoft MCP Gateway:Kubernetes 上的 MCP 数据面与控制面ai/sources
wiki/ai/sources/microsoft-mcp-gateway-kubernetes-control-plane.md
Mill — guardrailed local agentic workflowsai/sources
wiki/ai/sources/mill-guardrailed-local-agentic-workflows.md
MobileController:iOS Agent 的观察阶梯与像素真相ai/sources
wiki/ai/sources/mobilecontroller-ios-agent-observation-ladder.md
NVIDIA NeMo Relay — agent runtime control and trajectory layerai/sources
wiki/ai/sources/nemo-relay-agent-runtime-control.md
NeuroArxiv — prior-art scouting skillai/sources
wiki/ai/sources/neuroarxiv-prior-art-scouting-skill.md
Nothing New Under The Sun — research-first scouting skillai/sources
wiki/ai/sources/nothing-new-under-the-sun-research-first-scouting.md
OKF Gem — local Open Knowledge Format toolkit for agentsai/sources
wiki/ai/sources/okf-gem-local-knowledge-bundles.md
OKFy — purpose-shaped knowledge bundles for agentsai/sources
wiki/ai/sources/okfy-purpose-shaped-knowledge-bundles.md
OSWorld-MCP:计算机使用 agent 的 MCP 工具调用评测ai/sources
wiki/ai/sources/osworld-mcp-tool-invocation-computer-use-benchmark.md
OSWorld-V2:版本化 computer-use benchmark 与可复现环境合同ai/sources
wiki/ai/sources/osworld-v2-release-versioned-computer-use-benchmark.md
Open Science Skills — source-grounded research skill libraryai/sources
wiki/ai/sources/open-science-skills-research-skill-library.md
OpenBench — Harness Benchmarks for Coding Agentsai/sources
wiki/ai/sources/openbench-harness-benchmark.md
OpenBenchmark:从真实 agent 轨迹生成任务、rubric 与成本榜单ai/sources
wiki/ai/sources/openbenchmark-trajectory-to-benchmark.md
OpenSpec + Superpowers + gstack:一套让 AI 从「写代码」到「做项目」的组合拳ai/sources
wiki/ai/sources/OpenSpec+Superpowers+gstack-让AI从写代码到做项目.md
OptMem — permanent append-only memory for AI agentsai/sources
wiki/ai/sources/optmem-permanent-agent-memory.md
PERFOPT-Bench — Evaluating Coding Agents on Software Performance Optimizationai/sources
wiki/ai/sources/perfopt-bench-performance-optimization-agents.md
PIBench:支付集成场景的 coding-agent benchmarkai/sources
wiki/ai/sources/pibench-payment-integration-benchmark.md
PM Manager — local project governance skill packai/sources
wiki/ai/sources/pm-manager-local-project-governance.md
PawBench:Model × Harness 共评测的 Agent Benchmarkai/sources
wiki/ai/sources/pawbench-model-harness-coevaluation.md
Performance Skill:面向 .NET 性能工程的 agent skillai/sources
wiki/ai/sources/performance-skill-dotnet-performance-agent-skill.md
Permit MCP Gateway — enterprise trust layer for MCP agentsai/sources
wiki/ai/sources/permit-mcp-gateway-enterprise-trust-layer.md
Portcullis — Claude Code security hooksai/sources
wiki/ai/sources/portcullis-claude-code-security-hooks.md
Prismor — agent runtime security control planeai/sources
wiki/ai/sources/prismor-runtime-control-plane.md
Proctor — signed benchmark integrity bundlesai/sources
wiki/ai/sources/proctor-signed-benchmark-integrity-bundles.md
Promtact — inline policy enforcement point for agent tool callsai/sources
wiki/ai/sources/promtact-inline-policy-enforcement-point.md
Proof-or-Stop:证据门控的 agent 生命周期控制ai/sources
wiki/ai/sources/proof-or-stop-evidence-gated-lifecycle-control.md
Prospective multi-pathogen disease forecasting usiai/sources
wiki/ai/sources/Prospective_multi-pathogen_disease_forecasting_usi.md
ProveKit MCP — red-team-hardened MCP serverai/sources
wiki/ai/sources/provekit-mcp-redteam-hardened-mcp-server.md
REDAgentBench:以服务状态验证危害,而不是以最终拒绝判断安全ai/sources
wiki/ai/sources/redagentbench-state-grounded-safety-measurement.md
RSIHub:Frozen Evaluator 驱动的 Agent 自我改进工作区ai/sources
wiki/ai/sources/rsihub-frozen-evaluator-agent-self-improvement.md
RealReplicaBench — stateful business agent benchmarkai/sources
wiki/ai/sources/realreplicabench-stateful-business-agent-benchmark.md
RepoAgentBench:把真实合并 PR 变成私有 coding-agent benchmarkai/sources
wiki/ai/sources/repoagentbench-private-pr-coding-agent-benchmark.md
RepoPrompt CE — native context engineering app for coding agentsai/sources
wiki/ai/sources/repoprompt-ce-context-engineering-app.md
RepoTrials — private repository coding-agent evaluationsai/sources
wiki/ai/sources/repotrials-private-repository-coding-agent-eval.md
RimZ — terminal-native control room for coding agent fleetsai/sources
wiki/ai/sources/rimz-agent-fleet-control-room.md
SAIL Skill — Secure AI Lifecycle as an agent skillai/sources
wiki/ai/sources/sail-skill-secure-ai-lifecycle.md
SETA — Scaling Environments for Terminal Agentsai/sources
wiki/ai/sources/seta-scaling-environments-terminal-agents.md
SWE Refactor Bench:长程全仓迁移的三阶段评测ai/sources
wiki/ai/sources/swe-refactor-bench-whole-repository-migration.md
SWE-Marathon:超长程软件工程 agent benchmarkai/sources
wiki/ai/sources/swe-marathon-ultra-long-horizon-agent-benchmark.md
SWE-Review — closing coding-agent loops with agentic code reviewai/sources
wiki/ai/sources/swe-review-agentic-code-review.md
SWE-Touch:用户中途改动代码时的 coding-agent benchmarkai/sources
wiki/ai/sources/swe-touch-user-intervention-coding-benchmark.md
Sequoia Ascent 2026 — Software 3.0, Agentic Engineering, and Jagged Intelligenceai/sources
wiki/ai/sources/karpathy-sequoia-ascent-2026.md
Set-shifting Behavioral Test for Harnessed Agentsai/sources
wiki/ai/sources/set-shifting-behavioral-test-harnessed-agents.md
SimpleEnglish — controlled-language agent skill for unambiguous documentationai/sources
wiki/ai/sources/simpleenglish-controlled-language-agent-skill.md
Sitegeist — visual diversity benchmark for coding agentsai/sources
wiki/ai/sources/sitegeist-visual-diversity-benchmark.md
Sitegeist:检测 coding agents 的视觉设计趋同ai/sources
wiki/ai/sources/sitegeist-visual-convergence-benchmark.md
Skill Bill — governed agent skill runtimeai/sources
wiki/ai/sources/skill-bill-governed-agent-skill-runtime.md
SkillForge — local-first agent skill runtime and registryai/sources
wiki/ai/sources/skillforge-local-first-agent-skill-runtime.md
Skills Are Not Islands — Agent Skill Supply Chainsai/sources
wiki/ai/sources/agent-skill-supply-chains.md
Smithers — durable observable agent workflowsai/sources
wiki/ai/sources/smithers-durable-agent-workflows.md
Stop Means Stop:agent framework 控制原语的执行缺口ai/sources
wiki/ai/sources/stop-means-stop-control-primitives.md
Suede Creator Skills — supervised open agent skill packai/sources
wiki/ai/sources/suede-creator-skills.md
The Evolution of the Agent Harness:从模型外壳到注意力接口ai/sources
wiki/ai/sources/attention-interface-agent-harness-evolution.md
The Unreasonable Effectiveness of RNNsai/sources
wiki/ai/sources/The_Unreasonable_Effectiveness_of_RNNs.md
Token Warden — benchmark-gated agent memoryai/sources
wiki/ai/sources/token-warden-benchmark-gated-agent-memory.md
TraceProbe — coding agent trajectory diagnosticsai/sources
wiki/ai/sources/traceprobe-trajectory-structure-diagnostics.md
Truco-Bench — covert signaling in LLM agentsai/sources
wiki/ai/sources/truco-bench-covert-agent-signaling.md
Vigiles — audit/lint/test/eval for agent harnessesai/sources
wiki/ai/sources/vigiles-agent-harness-audit.md
Vinv — runtime context bandits for coding agentsai/sources
wiki/ai/sources/vinv-runtime-context-bandits.md
WeaveBench:GUI+CLI 混合长程 computer-use 评测ai/sources
wiki/ai/sources/weavebench-hybrid-gui-cli-agent-benchmark.md
What to Keep, What to Forget — A Rate-Distortion View of Memory Compaction in LLMs and Agentsai/sources
wiki/ai/sources/memory-compaction-rate-distortion-agents.md
Workflow as Knowledge — Semantic Persistence for LLM Workflowsai/sources
wiki/ai/sources/workflow-as-knowledge-semantic-persistence.md
XORCISE:以 OpenTelemetry evidence 评分的 cyber-agent 任务环境ai/sources
wiki/ai/sources/xorcise-cyber-agent-evidence-benchmark.md
a2a-query:把 A2A task lifecycle 变成应用可消费的 reactive harnessai/sources
wiki/ai/sources/a2a-query-task-handle-approval-broker.md
agent-desktop — accessibility-tree computer-use runtimeai/sources
wiki/ai/sources/agent-desktop-accessibility-tree-computer-use-runtime.md
agent-egress-bench — security-tool corpus for AI agent egressai/sources
wiki/ai/sources/agent-egress-bench-security-tool-corpus.md
agent-session-io — harness-neutral session substrateai/sources
wiki/ai/sources/agent-session-io-harness-neutral-session-substrate.md
agent-vision-toolkit:给文本 Agent 加“视觉 harness”ai/sources
wiki/ai/sources/agent-vision-toolkit-vision-harness.md
agentic-community MCP Gateway & Registry:AI asset control planeai/sources
wiki/ai/sources/agentic-community-mcp-gateway-registry-ai-asset-control-plane.md
alint — model-backed lint rules for agent-generated codeai/sources
wiki/ai/sources/alint-model-backed-code-analysis.md
book-to-skill — technical books as agent skillsai/sources
wiki/ai/sources/book-to-skill-technical-book-agent-skill.md
claim-trace — measured-evidence claim atomai/sources
wiki/ai/sources/claim-trace-evidence-claim-atom.md
claude-scaffold:fork-and-fill 的 Claude Code 项目脚手架ai/sources
wiki/ai/sources/claude-scaffold-fork-and-fill-agent-project-bootstrap.md
coder_eval — evaluate AI coding agents & their skillsai/sources
wiki/ai/sources/coder-eval-skill-evaluation-ci.md
da-verify:程序化验证优于同模型自检的 agent harnessai/sources
wiki/ai/sources/da-verify-programmatic-verification-harness.md
deco Studio — MCP control plane for organizational agentsai/sources
wiki/ai/sources/deco-studio-mcp-control-plane.md
design-harness — evidence board for defensible designai/sources
wiki/ai/sources/design-harness-evidence-board.md
did-it claim-evidence reconciliationai/sources
wiki/ai/sources/did-it-claim-evidence-reconciliation.md
doctrine — Markdown information governance for LLM developmentai/sources
wiki/ai/sources/doctrine-markdown-information-governance.md
dsh-handbook:DeepSeek Harness 手册与生态观察ai/sources
wiki/ai/sources/dsh-handbook-ecosystem-observation.md
dsh-mini:便携 agent runtime 与低耦合 harness 拼装ai/sources
wiki/ai/sources/dsh-mini-portable-agent-runtime.md
dsh-nacos-bridge:Nacos AI Registry 到 dsh harness 的薄桥接层ai/sources
wiki/ai/sources/dsh-nacos-bridge-registry-to-harness-runtime.md
dsh-plugin-dev:把 DeepSeek Harness 插件开发压缩成 Agent Skillai/sources
wiki/ai/sources/dsh-plugin-dev-agent-skill.md
goalflow — Graph-Orchestrated Agent Loopai/sources
wiki/ai/sources/goal-flow-graph-orchestrated-agent-loop.md
halu-core — claim-grounded execution and reporting honesty benchmark engineai/sources
wiki/ai/sources/halu-core-claim-grounded-agent-evaluation.md
hermes-skill-loop — closed skill learning loopai/sources
wiki/ai/sources/hermes-skill-loop-closed-skill-learning-loop.md
llm-wiki 自我优化 2026-09-07:registry-to-runtime 与 A2A task-state gateai/sources
wiki/ai/sources/llm-wiki-optimization-2026-09-07.md
llm-wiki 自我优化 2026-09-08:asset-control-plane 与 marathon-eval receiptai/sources
wiki/ai/sources/llm-wiki-optimization-2026-09-08.md
llm-wiki 自我优化 2026-09-09:tool-route、action-interception 与 protocol-binding receiptai/sources
wiki/ai/sources/llm-wiki-optimization-2026-09-09.md
llm-wiki 自我优化 2026-09-10:验证实际边界,不再堆叠“支持”和“已通过”ai/sources
wiki/ai/sources/llm-wiki-optimization-2026-09-10.md
llm-wiki 自我优化 2026-09-11:可回读、可相信、获授权不是一回事ai/sources
wiki/ai/sources/llm-wiki-optimization-2026-09-11.md
llm-wiki 自我优化 2026-09-12:实现不覆盖标准,高分不覆盖弃权ai/sources
wiki/ai/sources/llm-wiki-optimization-2026-09-12.md
loop-board — autonomous PR task board with constraintsai/sources
wiki/ai/sources/loop-board-autonomous-pr-task-board.md
mcp-gauntlet:面向 MCP server 的 agentic 评估 harnessai/sources
wiki/ai/sources/mcp-gauntlet-agentic-mcp-server-evaluator.md
octopus-skill — host-agnostic long-horizon agent disciplineai/sources
wiki/ai/sources/octopus-skill-long-horizon-agent-discipline.md
pi-rlm — persistent single execute tool for agentsai/sources
wiki/ai/sources/pi-rlm-persistent-code-execution-tool.md
quorum — MCP agent collaboration spineai/sources
wiki/ai/sources/quorum-mcp-agent-collaboration-spine.md
redstamp:确定性 agent tool-call firewallai/sources
wiki/ai/sources/redstamp-deterministic-agent-firewall.md
reliable-cua:Computer-Use Agent Benchmark 的统计可靠性层ai/sources
wiki/ai/sources/reliable-cua-statistical-computer-use-eval.md
token-diet — always-on token-efficiency skill for coding agentsai/sources
wiki/ai/sources/token-diet-token-efficiency-skill.md
x-clipper 是一个浏览器扩展工具,通过 claude -p headless 加 tweetai/sources
wiki/ai/sources/x-clipper_是一个浏览器扩展工具通过_claude_-p_headless_加_tweet.md
企业 MCP 鉴权网关:persona × credential 与服务账号授权边界ai/sources
wiki/ai/sources/enterprise-mcp-persona-credential-boundaries.md
通过累积行为规则让 Coding Agent 跨会话自我改进ai/sources
wiki/ai/sources/self-improving-coding-agents-behavioral-rules.md
X / Karpathy Radar Daily Ingest Templateai/templates
wiki/ai/templates/x-karpathy-radar-daily-ingest.md