hermes-agent

Author	SHA1	Message	Date
Ben Barclay	4f7fe9bcff	fix(dashboard): surface Docker update guidance instead of generic failure (#34347 ) (#37085 ) The dashboard Update button's backend guard (#36263) already returns a structured {ok:false, error:"docker_update_unsupported", message, update_command} envelope (HTTP 200) when running in a Docker install, instead of surfacing a raw SystemExit. But the frontend ignored that envelope: runAction() only branched on a thrown error, so the 200 fell through to the action-status poll, which reported a generic "Action failed (exit 1)" toast and never showed the actual guidance. Now runAction() inspects the update response and, on the docker_update_unsupported case, surfaces the backend's guidance message plus the recommended re-pull command directly (success-styled, since it's actionable guidance — not a crash) without starting the poll. Closes #34347.	2026-06-02 10:36:10 +10:00
firefly	3a8d643d37	chore(release): map caojiguang@gmail.com in AUTHOR_MAP The fix commit preserves @caojiguang's authorship (from #31853); the release-notes AUTHOR_MAP gate requires their email to map to a GitHub username.	2026-06-01 17:31:40 -07:00
firefly	765790a216	test(weixin): regression suite for _api_post/_api_get timeout migration	2026-06-01 17:31:40 -07:00
Cao Jiguang	566669013f	fix(weixin): replace aiohttp ClientTimeout with asyncio.wait_for in _api_post/_api_get Cron delivery to WeChat fails with 'Timeout context manager should be used inside a task' because _api_post and _api_get use aiohttp's ClientTimeout directly. When the cron scheduler calls send() via asyncio.run_coroutine_threadsafe(), aiohttp cannot find a running task and raises RuntimeError. _upload_media, _download_bytes, and _download_remote_media already use asyncio.wait_for() to avoid this. Apply the same pattern to _api_post and _api_get — the two remaining iLink API helpers that still use the raw ClientTimeout approach. This fixes cron delivery errors seen on the WeChat platform adapter when meyo-external cron jobs attempt to deliver output to WeChat.	2026-06-01 17:31:40 -07:00
firefly	a1f76ba7e9	fix(gateway): recover extract-stripped tool responses on all platforms (#29346 ) The extract pipeline (extract_media/extract_images/extract_local_files + directive strips) can reduce a non-empty tool-using response to empty text_content with no deliverable attachment. The 'if text_content' send guard then silently skips delivery: a 'response ready' log with no 'Sending response', no error, and the answer never reaches the user. - A2: snapshot the pre-extract response; when extraction yields empty text and no image/local/media attachment, deliver the recovered original from the post-extract_media body (so a spaced MEDIA path can't leak). Applies on ALL platforms (supersedes the Discord-only #33842 and the unsafe raw-fallback #29499). - A3: loud delivery invariant - a non-empty response that produces nothing deliverable logs response_delivery_dropped at ERROR; every recovery logs response_delivery_recovered. No silent drop survives. - Factor a _strip_media_directives helper for the [[...]] strips; MEDIA stripping stays owned by extract_media, whose grammar handles spaced and quoted paths. - Salvaged + de-scoped the #33842 test harness to all platforms; added unrecoverable-drop and no-leak regression tests.	2026-06-01 17:31:32 -07:00
firefly	8bf498c21d	fix(gateway): scope final-delivery flags to turn-final segment (#29346 ) A streamed preamble ("Let me search...") finalized at a tool boundary routed through _try_fresh_final, which unconditionally set _final_response_sent=True even though it is a NON-final segment. The gateway then reads that flag as "final delivered" and suppresses the genuine final answer produced on the next API call, so the user silently gets nothing. Only reproduces with fresh_final_after_seconds > 0. - _try_fresh_final / _send_or_edit take is_turn_final; the segment-break call site passes is_turn_final=got_done so only the turn-final answer marks final-delivered. - _reset_segment_state clears the final-delivery flags at every tool boundary as defense-in-depth against any future premature setter. - Failing-first regression + happy-path no-duplicate test.	2026-06-01 17:31:32 -07:00
Teknium	92273e4f57	docs: add 25 new community user stories to the collage (#37048 ) Sourced from X/Twitter, blogs (Medium/Substack/dev.to), and YouTube since the last refresh. Deduped against the existing 237 entries by id, url, and author. 237 -> 262 stories. Highlights: 24/7 Mac Mini agent at $21/mo (@witcheer), automated TikTok slideshow factory (@cyrilXBT), per-client isolated profiles as an AI-ops business (@IBuzovskyi), PM briefing 20->8min (@aakashgupta), Railway+Telegram deploy gotchas (Tessa Kriesel), compounding-cost field report (chintanonweb), 18-agent Kanban fleet (Tonbi), and several daily-automation setups.	2026-06-01 17:01:18 -07:00
kshitijk4poor	0fdab53ef0	feat(cli): ranked fuzzy search in the curses model picker Wires the salvaged search helpers into the shared curses menu driver and turns on type-to-filter for the CLI model pickers (the 100+ model lists that previously required scrolling). - Search lives in the shared `_run_curses_menu` driver behind a `searchable` flag + `search_labels`, so both `curses_radiolist` and `curses_single_select` get it without per-menu duplication. `/` opens the filter, BACKSPACE edits, Ctrl+U clears, ESC clears the filter then cancels. Returned values are always original item indices. - `_filter_indices` RANKS matches (best-first) via a Python port of the TS scorer in ui-tui/src/lib/fuzzy.ts and web/src/lib/fuzzy.ts. The port is byte-identical in score: same per-char bonuses, prefix (+8) and exact (+20) bonuses, camelCase/word-boundary detection (matching on the lowercased target, boundary on the original case), and the -len*0.01 length tiebreak — so the CLI, TUI, and WebUI rank results identically. A cross-language parity test pins the exact scores. - `_prompt_model_selection` (the canonical picker across the model flows) and the custom-provider model list pass `searchable=True`. - Split `_decode_menu_key` out of `read_menu_key` so the search loop can peek the raw key (catch `/`) before nav decoding. - ESC during active search now clears the query (restores the full list) so a no-match filter can't strand the user; printable-key capture is restricted to ASCII to avoid Latin-1 mojibake. - Update two setup-menu tests whose mock signatures predate the new `searchable` kwarg; add ranked-scorer + parity + state-machine tests.	2026-06-01 16:58:58 -07:00
Harish Kukreja	53f598e7a2	feat(cli): add fuzzy search helpers for curses pickers Pure, refactor-independent helpers for type-to-filter search in the curses single-/radio-select menus: subsequence matching, filtered-index mapping, cursor reconciliation, scroll clamping, and an active-search key handler, plus unit tests. Salvaged from #22758 (the curses event loop was since refactored into a shared driver on main, so the integration is rebuilt in a follow-up commit; these pure helpers and their tests carry over unchanged).	2026-06-01 16:58:58 -07:00
kshitijk4poor	7527e7aeac	feat: fuzzy search for the model picker (WebUI + TUI) Adds fuzzy subsequence matching with quality ranking to the model pickers, replacing the WebUI's exact-substring filter and giving the TUI a search where it previously had none. - New fuzzy scorer (ui-tui/src/lib/fuzzy.ts + an identical copy at web/src/lib/fuzzy.ts, since the two are separate TS packages with no shared module). Matches a query as an ordered subsequence (so `g4o` matches `gpt-4o`), scores by quality (exact > prefix > word-boundary > contiguous > scattered) and returns matched character positions for highlighting. Multi-token AND semantics (`clad snnt` -> claude-sonnet). 15 vitest tests cover the algorithm. - WebUI ModelPickerDialog: ranked fuzzy filter on providers + models; matched characters in model rows are highlighted via <mark>. - TUI modelPicker: type-to-filter on the provider and model stages with live ranking. Backspace edits the filter, Ctrl+U clears it, Esc clears a non-empty filter before navigating back. Persist-global / disconnect shortcuts moved from g/d to Ctrl+G / Ctrl+D so letters feed the filter. Closes #30849	2026-06-01 16:58:58 -07:00
Teknium	c45593ceae	docs: expand quickstart Skills section (#37047 ) * fix(file_tools): block agent writes to ~/.hermes/config.yaml to prevent silent approval bypass * fix(approval): pair terminal-side gate for ~/.hermes/config.yaml writes Subway2023's #14639 blocks write_file/patch to ~/.hermes/config.yaml, but the terminal side was only partially paired: echo>/tee/cp/mv to config.yaml already tripped the project-config pattern, while `sed -i` and direct edits slipped through with auto-approve. An unpaired write_file deny is theater per SECURITY.md — the agent could flip approvals.mode=off via `sed -i` and the mtime-keyed config cache reloads it mid-session. config.yaml IS the security policy (approvals.mode/yolo/permanent allowlist live there), so it warrants real pairing, not a half-door. Add a _HERMES_CONFIG_PATH fragment mirroring _HERMES_ENV_PATH, fold it into _SENSITIVE_WRITE_TARGET (covers tee/>/>>/cp/mv), and add sed -i coverage for both config.yaml and .env. Pins 9 regression tests including no-regression guards (reads pass, /tmp writes pass). Co-authored-by: sbw2025 <subw3@mail2.sysu.edu.cn> * chore(release): map Subway2023 for PR #14639 salvage * docs: expand quickstart Skills section The Skills section was two bare commands with no framing — it never said what a skill is, how skills load, or what the install slug means. Expanded to explain the concept, the bundled catalog, install/browse/use flow, and slash-command activation. Removed the inaccurate /skills chat-command hint (skills become individual /<name> commands; hermes skills is the CLI verb). --------- Co-authored-by: sbw2025 <subw3@mail2.sysu.edu.cn>	2026-06-01 16:56:50 -07:00
firefly	128da68823	test(tools): characterize tool-surface TERMINAL_CWD contract (#29265 ) Port PR #29365's tool-surface contract test: terminal/file/execute_code already honor TERMINAL_CWD (out of scope for the resolver cluster). Pinning the behavior makes the supersession of #29365 airtight and guards against a future refactor silently regressing the workspace contract.	2026-06-01 16:55:04 -07:00
firefly	ac0cce5f3f	test(agent): pin whitespace-strip and OSError-propagation in runtime_cwd Cover the two new hardening behaviors that were unpinned: whitespace-only TERMINAL_CWD falling through to getcwd/None, and OSError from the getcwd fallback arm propagating to the build_environment_hints try/except guard.	2026-06-01 16:55:04 -07:00
firefly	75f478750c	docs(test): correct None-semantics comment in test_runtime_cwd (discovery not skipped)	2026-06-01 16:55:04 -07:00
firefly	eadfeef60e	docs(agent): correct resolve_context_cwd comment (None → caller getcwd fallback, not skip)	2026-06-01 16:55:04 -07:00
firefly	f90777a6b8	refactor(prompt): route context-file cwd through runtime_cwd resolver	2026-06-01 16:55:04 -07:00
firefly	c79b80a8a5	test(prompt): place cwd regression tests in TestEnvironmentHints (drop redundant docker case)	2026-06-01 16:55:04 -07:00
firefly	16047655b5	fix(prompt): show configured working directory in system prompt (closes #24882 , #24969 , #27383 , #29265 )	2026-06-01 16:55:04 -07:00
firefly	2564760d7a	test(agent): pin context_cwd isdir-skip asymmetry and tilde expansion	2026-06-01 16:55:04 -07:00
firefly	4bc7296042	feat(agent): add runtime_cwd resolver (single source of truth for working dir)	2026-06-01 16:55:04 -07:00
teknium1	f1237aa95b	chore(release): map maxcz79 author email for AUTHOR_MAP	2026-06-01 16:36:43 -07:00
maxcz79	32032e1e2d	fix(simplex): avoid reconnecting healthy idle websocket Do not treat lack of application-level SimpleX events as a stale WebSocket. The websockets client already uses protocol ping/pong for connection liveness, so quiet but healthy connections should not be closed by the health monitor.	2026-06-01 16:36:43 -07:00
Teknium	e946f49ab5	fix(models): add gemini-3.5-flash to Gemini OAuth + API-key pickers (#37046 ) * fix(file_tools): block agent writes to ~/.hermes/config.yaml to prevent silent approval bypass * fix(approval): pair terminal-side gate for ~/.hermes/config.yaml writes Subway2023's #14639 blocks write_file/patch to ~/.hermes/config.yaml, but the terminal side was only partially paired: echo>/tee/cp/mv to config.yaml already tripped the project-config pattern, while `sed -i` and direct edits slipped through with auto-approve. An unpaired write_file deny is theater per SECURITY.md — the agent could flip approvals.mode=off via `sed -i` and the mtime-keyed config cache reloads it mid-session. config.yaml IS the security policy (approvals.mode/yolo/permanent allowlist live there), so it warrants real pairing, not a half-door. Add a _HERMES_CONFIG_PATH fragment mirroring _HERMES_ENV_PATH, fold it into _SENSITIVE_WRITE_TARGET (covers tee/>/>>/cp/mv), and add sed -i coverage for both config.yaml and .env. Pins 9 regression tests including no-regression guards (reads pass, /tmp writes pass). Co-authored-by: sbw2025 <subw3@mail2.sysu.edu.cn> * chore(release): map Subway2023 for PR #14639 salvage * fix(models): add gemini-3.5-flash to Gemini OAuth + API-key pickers #34581 swapped gemini-3-flash-preview -> gemini-3.5-flash in the OpenRouter and Nous lists but missed the curated Gemini catalogs, so the Google OAuth (google-gemini-cli) picker still offered the retired gemini-3-flash-preview slug and gemini-3.5-flash was unselectable. Per Google's docs gemini-3-flash-preview was renamed to gemini-3.5-flash and is served via Cloud Code Assist, so this completes the rename for: - google-gemini-cli (OAuth/Code Assist) picker - gemini (API-key) picker - gemini provider default_aux_model copilot keeps gemini-3-flash-preview (separate backend, own slug). --------- Co-authored-by: sbw2025 <subw3@mail2.sysu.edu.cn>	2026-06-01 16:31:13 -07:00
Teknium	1ffa22ee6b	fix(minimax): drop stale ≤204,800 cache entries for MiniMax-M3 (#36726 ) M3 is 1M context, but pre-catalog builds resolved it via the generic 'minimax' catch-all (204,800) and persisted that to the context-length cache. Step 1 of get_model_context_length returned the cached value directly before reaching the 'minimax-m3' (1M) catalog entry, so users who first probed M3 on an older build were stuck at 204K forever (e.g. /new in the Telegram gateway showing 'Context: 204K tokens (detected)'). Mirror the existing Kimi/Codex stale-cache guards: when a cached entry for a minimax-m3 slug is <= 204,800, drop it and re-resolve. M2.x slugs (correctly 204,800) are untouched since they don't match the M3 name.	2026-06-01 14:59:07 -07:00
Ben	b9646276fd	fix(utils): guard os.fchmod for Windows in atomic_json_write os.fchmod is Unix-only; the Windows os module has no fchmod (only chmod). Passing mode= (e.g. 0o600 when saving the Hindsight config during `hermes memory setup`) crashed on Windows with: AttributeError: module 'os' has no attribute 'fchmod' Guard the fchmod fast-path with hasattr(os, "fchmod"). Skipping it on Windows is safe: mkstemp already creates the temp file as 0o600, and the existing post-replace os.chmod(real_path, mode) — already wrapped in try/except — applies the final mode durably (as far as Windows honors it). Adds regression tests: one simulating a Windows os module without fchmod (must not raise), and one asserting the durable 0o600 mode on POSIX. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>	2026-06-01 09:57:10 -07:00
kshitij	a5371b3e68	chore: add benfrank241 to AUTHOR_MAP (#36898 ) Maps ben.bartholomew@vectorize.io -> benfrank241 so the contributor attribution audit passes when their commit lands via #36824.	2026-06-01 16:47:07 +00:00
teknium1	ef3a650f05	chore(release): map Subway2023 for PR #14639 salvage	2026-06-01 03:29:48 -07:00
teknium1	4e9d886d9d	fix(approval): pair terminal-side gate for ~/.hermes/config.yaml writes Subway2023's #14639 blocks write_file/patch to ~/.hermes/config.yaml, but the terminal side was only partially paired: echo>/tee/cp/mv to config.yaml already tripped the project-config pattern, while `sed -i` and direct edits slipped through with auto-approve. An unpaired write_file deny is theater per SECURITY.md — the agent could flip approvals.mode=off via `sed -i` and the mtime-keyed config cache reloads it mid-session. config.yaml IS the security policy (approvals.mode/yolo/permanent allowlist live there), so it warrants real pairing, not a half-door. Add a _HERMES_CONFIG_PATH fragment mirroring _HERMES_ENV_PATH, fold it into _SENSITIVE_WRITE_TARGET (covers tee/>/>>/cp/mv), and add sed -i coverage for both config.yaml and .env. Pins 9 regression tests including no-regression guards (reads pass, /tmp writes pass). Co-authored-by: sbw2025 <subw3@mail2.sysu.edu.cn>	2026-06-01 03:29:48 -07:00
sbw2025	8f2931e3ee	fix(file_tools): block agent writes to ~/.hermes/config.yaml to prevent silent approval bypass	2026-06-01 03:29:48 -07:00
Teknium	023149f665	fix(agent): stop reporting broken streams as output-length truncation (#36705 ) A stream that drops mid-response after tokens are delivered (peer-closed connection, stale-stream reconnect) is converted into a synthetic finish_reason="length" stub. The conversation loop treated that network stall as a max-output-tokens truncation: when the dropped content was a tool call it retried exactly once, then hard-failed with "Response truncated due to output length limit" — even on large-output models that never hit any cap (e.g. Opus). - Tool-call truncation now retries up to 3 times (was 1) with a progressive max_tokens boost, and is stub-aware: a PARTIAL_STREAM_STUB_ID stall prints "Stream interrupted mid tool-call — retrying (n/3)" instead of the false "model hit max output tokens", and the give-up message distinguishes a network drop from a real truncation. - Length-continuation retries preserve the original request's output cap as a floor, so a high provider/model default isn't silently downshifted to 8K/12K on retry. - Added _requested_output_cap_from_api_kwargs() helper. Tests: stub-stall mid-tool-call recovery within 3 retries; continuation preserves a large provider-default output cap. Fixes #26425. Salvages the substance of #26427 (cap floor) and #9525 (retry bump), adapted to the post-refactor conversation_loop.py which handles all three api_modes uniformly. Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com> Co-authored-by: ygd58 <ygd58@users.noreply.github.com>	2026-06-01 03:01:20 -07:00
Teknium	b571ec298d	feat(dashboard): full administration panel — MCP, pairing, webhooks, credentials, memory, gateway, ops (#36704 ) * feat(dashboard): backend API for MCP, pairing, webhooks, credential pool, memory, gateway lifecycle Adds REST endpoints so a remote admin can manage these without CLI access: - MCP servers: list/add/remove/test (config.yaml parity with hermes mcp) - Pairing: list/approve/revoke/clear-pending messaging codes - Webhooks: list/subscribe/remove (hot-reloaded JSON store) - Credential pool: list/add/remove rotation keys (via CredentialPool API) - Memory provider: status/select/disable/reset - Gateway lifecycle: start/stop (restart+update already existed) Secrets redacted on read; usable values only reach the agent at session start. All endpoints sit behind the existing dashboard auth gate. * feat(dashboard): backend API for ops + skills hub - Ops actions (spawned, log-tailed via /api/actions): doctor, security audit, backup, import, checkpoints prune - Ops reads (structured JSON): hooks list + allowlist status, checkpoints list with per-session size - Skills hub actions (spawned): install / uninstall / update - Registers new action log files for all spawn-based endpoints All gated by the existing dashboard auth middleware. * feat(dashboard): admin pages for MCP, pairing, webhooks, and system ops Adds four new dashboard pages + nav entries so a remote admin can manage Hermes without CLI access: - MCP: list/add/remove/test MCP servers - Webhooks: list/create/delete subscriptions (one-time secret reveal) - Pairing: approve/revoke/clear messaging pairing codes - System: gateway start/stop/restart, memory provider + reset, credential pool add/remove, ops (doctor/audit/backup/import/skills update) with a live action-log viewer, checkpoints prune, shell-hooks status api.ts: client methods + types for all new endpoints. App.tsx: routes + sidebar nav (plain labels, no i18n key required). Verified: tsc -b clean, production build succeeds, new pages lint clean, zero new eslint errors in App.tsx. * test(dashboard): cover admin API endpoints 20 tests across MCP, credential pool, memory, pairing, webhooks, ops, plus an auth-gate parametrize that asserts every admin endpoint requires the session token. Asserts request contract + CLI-config parity, not catalog values (per the no-change-detector-tests rule). * docs(dashboard): document MCP, Webhooks, Pairing, and System admin pages Adds Pages sections for the four new admin tabs and an Admin-endpoints table to the REST API reference. Updates the page description to reflect the dashboard's expanded role as a full administration panel.	2026-06-01 02:58:02 -07:00
Teknium	2ed96372ad	feat(skills): blank-slate skills — install --no-skills + opt-out/opt-in (#36228 ) * feat(install): --no-skills flag for blank-slate default profile Add an install-time --no-skills flag so the default ~/.hermes profile can be created with zero bundled skills, matching what `hermes profile create --no-skills` already does for named profiles. The flag writes $HERMES_HOME/.no-bundled-skills and skips the install-time seed. sync_skills() now honors that marker with an early return (skipped_opt_out=True), so neither the installer, a later `hermes update`, nor a direct sync re-injects bundled skills into a profile that opted out. Previously the marker was only checked by seed_profile_skills() (named profiles); the default profile had no opt-out and `hermes update` would re-seed it every time. Tests: TestNoBundledSkillsOptOut covers marker-present (no-op) and marker-absent (normal seed) paths. * feat(skills): hermes skills opt-out / opt-in for existing profiles Adds an interactive counterpart to the install-time --no-skills flag so an already-installed profile (default or named) can toggle the .no-bundled-skills marker without reinstalling. - `hermes skills opt-out` writes the marker (stop future seeding). Safe by default: nothing on disk is touched. - `hermes skills opt-out --remove` ALSO deletes already-present bundled skills, but ONLY ones that are manifest-tracked AND byte-identical to their origin hash. User-edited bundled skills, hub-installed skills, and hand-written skills are never removed. Previews + confirms before deleting (--yes to skip). - `hermes skills opt-in [--sync]` removes the marker and optionally re-seeds immediately. Core logic lives in tools/skills_sync.py (set_bundled_skills_opt_out, is_bundled_skills_opt_out, remove_pristine_bundled_skills) reusing the existing manifest origin-hash machinery for the safety check. Tests: TestOptOutToggleAndRemove covers marker toggle idempotency and proves user-modified + non-bundled skills survive --remove. * docs: blank-slate skills — install --no-skills + opt-out/opt-in - features/skills.md: new 'Starting with a blank slate' section covering the install flag, profile-create flag, and runtime opt-out/opt-in, with a safe-by-default note. - reference/cli-commands.md: document the new skills opt-out / opt-in subcommands + examples. - reference/profile-commands.md: fix the marker filename (was .no-skills, actually .no-bundled-skills) and cross-link the runtime commands. Validated with a full docusaurus build (exit 0); the three edited pages compile clean with no new warnings.	2026-06-01 02:57:57 -07:00
Teknium	70e1571d89	feat(curator): prune built-in skills after inactivity + track usage for all skills (#36701 ) Two related changes to the skill curator: 1. Built-in pruning. New curator.prune_builtins config (default on) lets the curator archive bundled built-in skills after the inactivity period, not just agent-created ones. A .curator_suppressed list tells the update-time re-seeder (tools/skills_sync) to leave pruned built-ins archived, so the prune is durable across `hermes update`. Built-ins are seeded with a baseline record on first sight, so the inactivity clock starts at upgrade time -- no mass-prune on the first run. Hub-installed skills are never pruned regardless of the flag. Restoring a built-in clears its suppression. 2. Usage tracking for all skills. Telemetry (view/use/patch) was wrongly gated behind curation-eligibility, so built-ins were tracked only when prunable and hub skills never. Telemetry is observability and is now decoupled from curation: every skill accrues usage counts regardless of provenance, while lifecycle mutators (set_state/set_pinned/mark_agent_created) stay curation-gated. New usage_report() + provenance() expose all skills with an agent/bundled/hub tag.	2026-06-01 02:07:32 -07:00
Teknium	0622a70eb4	feat(gateway): bring /undo [N] to messaging platforms (parity with CLI/TUI) (#36699 ) Gateway /undo was wired into every platform but still ran the old single-turn hard-truncate. Now it matches the CLI/TUI: /undo [N] backs up N user turns (default 1, clamps to oldest), soft-deletes the truncated rows on disk (active=0, kept for audit, hidden from re-prompts and search) via SessionDB.rewind_to_message, evicts the cached agent so the next turn rebuilds from the active-only transcript (the gateway's equivalent of the CLI's in-place history surgery + memory invalidation), and echoes the backed-up message text so the user can copy/edit and resend — platforms have no editable composer to prefill. - gateway/session.py: SessionStore.rewind_session(session_id, n) wraps the soft-delete primitive; load_transcript already returns active-only - gateway/run.py: _handle_undo_command parses [N], calls rewind_session, evicts the agent, echoes target text; confirm-prompt detail is count-aware - locales: undo.removed gains {turns}; new undo.invalid_count, all 16 langs - tests: tests/gateway/test_undo_rewind_session.py (6 cases)	2026-06-01 02:04:14 -07:00
Teknium	ba6ffd4ff1	fix(skills-guard): stop flagging benign skill content + honor skill ignore files (#36231 ) The skill security scanner blocked legitimate community skills on three intrinsic false-positive patterns: - read_secrets_file matched `cat > file.env <<` heredocs (writing the user's own keys into their own local .env), not just `cat file.env` reads. Exclude output redirections. - allowed-tools frontmatter is REQUIRED by the agent-skill spec; every compliant skill declares it. Drop from HIGH privilege_escalation to a LOW informational finding so it no longer drives the verdict. - python_os_environ flagged `os.environ.get("CONFIG_VAR")` config reads as HIGH exfiltration. Exempt non-secret `.get()` reads; add a dedicated CRITICAL python_environ_get_secret pattern so secret-named reads (OPENAI_API_KEY etc.) are still caught. Also: scan_skill() now honors a skill-provided .skillignore / .clawhubignore (gitignore-style) so dev/docs artifacts shipped in a skill root are excluded from both structural checks and pattern scanning. SKILL.md is never ignorable. 80 tests pass (64 existing + 16 new).	2026-06-01 01:58:48 -07:00
Teknium	9074a154c5	feat: explain Quick Setup vs Full setup inline in the first-time setup menu (#36227 ) The setup-mode chooser showed two bare labels ('Quick Setup (Nous Portal) — OAuth login, model & messaging' / 'Full setup — configure everything') that didn't explain what Quick Setup actually is. Expand both labels inline so each choice line carries a concise explanation: Quick Setup (Nous Portal) — free OAuth login, no API keys, model + tools Full setup — configure every provider, tool & option yourself (bring your own keys) Single-file change to the choice labels; no new plumbing.	2026-06-01 01:58:30 -07:00
Teknium	92a567db2d	fix(ci): regen model catalog + stop gui tests consuming macos-fixup subprocess calls (#36687 ) Two pre-existing failures on main, unrelated to each other: - test_model_catalog: website/static/api/model-catalog.json was stale vs _PROVIDER_MODELS — minimax/minimax-m2.7 was renamed to minimax/minimax-m3 without regenerating the committed manifest. Ran scripts/build_model_catalog.py. - test_gui_command: the macOS relaunchable-signing fixup (_desktop_macos_relaunchable_fixup) makes two subprocess.run calls (xattr + codesign) on darwin before launch. The two darwin GUI tests set sys.platform='darwin' and mock subprocess.run with a 2-element side_effect (pack + launch), so the fixup's calls drained the iterator -> StopIteration. Mock out the fixup in those two tests so the subprocess accounting stays focused on pack/launch.	2026-06-01 01:39:03 -07:00
Teknium	e1951ce704	fix(memory): only forward rewound kwarg when set The on_session_switch fan-out passed rewound=rewound unconditionally, injecting rewound=False into every provider's **kwargs on the common /resume, /branch, /new, and compression paths. Providers that capture extra kwargs into an 'extra' dict (and the exact-dict-equality tests guarding them) broke. Forward rewound only when truthy; /undo sets it explicitly, everyone else stays clean.	2026-06-01 01:22:38 -07:00
Teknium	3f7d1c801d	feat(undo): /undo [N] backs up N user turns with prefill + soft-delete Extends the existing /undo command from a single in-memory exchange removal into a full rewind: back up N user turns (default 1), soft-delete the truncated rows in SessionDB (active=0, kept for audit, hidden from re-prompts and search), notify memory providers, and prefill the composer with the backed-up message text for editing — CLI and TUI. Reuses the SessionDB rewind primitives, the on_session_switch(rewound=True) memory hook, and the TUI command.dispatch prefill payload from SaguaroDev's #21910 work, wired to /undo [N] instead of a separate /rewind picker. - cli.py: undo_last(n, prefill) — in-memory truncate + SQLite soft-delete + agent surgery (system-prompt invalidate, flush-index reset) + memory notify + editable buffer prefill; /undo dispatch parses optional count; checkpoint-rollback caller passes prefill=False - tui_gateway/server.py: command.dispatch undo branch (was rewind) parses count, picks Nth-from-last user turn, clamps to oldest - commands.py: /undo gains [N] args_hint - tests: rename + expand TUI suite (multi-turn, clamp, invalid-count) - release.py: AUTHOR_MAP entry for SaguaroDev Co-authored-by: SaguaroDev <74339271+SaguaroDev@users.noreply.github.com>	2026-06-01 01:22:38 -07:00
SaguaroDev	243e836dce	feat(tui): wire /rewind through command.dispatch + prefill payload (#21910 ) Adds the TUI half of the /rewind feature so the Ink terminal UI gets the same affordance as the prompt_toolkit CLI. Python side (tui_gateway/server.py): - /rewind added to _PENDING_INPUT_COMMANDS so slash.exec rejects it and the TUI falls through to command.dispatch (the only path with access to live session state + memory hooks). - New command.dispatch branch for name == "rewind": v1 auto-picks the most recent user turn (Claude-Code-style single- step undo), calls SessionDB.rewind_to_message, refreshes the in-memory history, fires _memory_manager.on_session_switch with rewound=True, and returns the new "prefill" payload. - A dedicated picker overlay (multi-step rewind) is tracked as a follow-up to #21910. TS side (ui-tui/src/): - New "prefill" variant on CommandDispatchResponse + asCommandDispatch validator. Mirrors "send" but does NOT auto-submit; the client drops the message into the composer for editing. - createSlashHandler renders the optional notice via sys() and calls ctx.composer.setInput(d.message), letting the user edit-and-resubmit the rewound turn — the core UX promised by the issue. Tests: - 7 new tui_gateway tests covering prefill payload shape, in-memory history truncation, DB soft-delete, memory-provider notification (rewound=True), busy-session refusal, missing-session error, and registry placement in _PENDING_INPUT_COMMANDS. - Extended asCommandDispatch vitest covering the new prefill variant (with + without notice, and rejection of malformed payloads). Out of scope for v1 (tracked as #21910 follow-up): - Dedicated picker overlay in Ink (the multi-step rewind UI). v1 auto- picks the most recent user turn, matching the most common case. - Gateway platforms (Telegram, Discord, etc.) — issue scopes v1 to CLI + TUI only.	2026-06-01 01:22:38 -07:00
SaguaroDev	31cfa08c66	feat(memory): add rewound kwarg to on_session_switch hook	2026-06-01 01:22:38 -07:00
SaguaroDev	3e59be0c41	feat(state): add messages.active flag + rewind primitives (#21910 ) Schema v12 adds: - messages.active (default 1) — soft-delete flag for /rewind - sessions.rewind_count (default 0) — audit counter - idx_messages_session_active deferred index New SessionDB methods: - rewind_to_message(session_id, target_message_id) — soft-deletes rows >= target_id, refuses non-user targets, increments rewind_count - restore_rewound(session_id, since_message_id) — undo for stretch goal - list_recent_user_messages — picker source Existing methods get include_inactive kwarg (default False): - get_messages, get_messages_as_conversation, search_messages. Rewound rows excluded from session_search by default — opt-in for audit. The deferred index pattern (DEFERRED_INDEX_SQL run after _reconcile_columns) avoids 'no such column: active' on legacy pre-v12 databases, since executescript(SCHEMA_SQL) runs before column reconciliation.	2026-06-01 01:22:38 -07:00
kshitijk4poor	6c73e8ffaa	fix(gateway): keep code blocks verbatim in cleaned text when media present Self-review of the code-block masking fix: the cleanup path ran media_pattern.sub('') over the _mask_protected_spans() copy of the text and assigned that back to 'cleaned', so whenever a real MEDIA: tag was delivered (if media: branch), every fenced code block / inline code / blockquote in the reply was blanked to whitespace in the user-visible text. Now mask only a length-equal copy of 'cleaned' to locate the real tag spans, then delete those spans from the unmasked 'cleaned' — masking is a locator, not a text rewrite. Protected spans survive verbatim. Strengthens the existing mixed-code test (it only asserted 'Done.' survived, not the code block) and adds an inline-code-survives regression test. Both fail on the old sub-based code and pass now.	2026-06-01 00:00:26 -07:00
kshitijk4poor	ec6261ae2f	chore(release): add VinciZhu to AUTHOR_MAP for #16721 salvage	2026-06-01 00:00:26 -07:00
liuhao1024	3ccf4fdc6d	fix(gateway): skip MEDIA: tags inside code blocks and blockquotes extract_media() scanned the full response text without distinguishing live delivery tags from example paths in fenced code blocks, inline code spans, and blockquotes. This caused false positives where the agent's explanation of MEDIA: syntax (or tool output containing example paths) was stripped from user-visible text and the path was added to the media delivery list. Added _mask_protected_spans() helper that replaces protected regions with equal-length whitespace before regex matching, preserving match offsets. The helper skips backtick-quoted paths in MEDIA: tags to maintain existing path extraction behavior. Fixes #35695	2026-06-01 00:00:26 -07:00
VinciZhu	521d06975e	fix(gateway): restrict auto-appended media to producer tools	2026-06-01 00:00:26 -07:00
kshitijk4poor	fb1b681b3b	fix(gateway): keep JSON-embedded MEDIA: text verbatim in cleaned output Self-review of #34375 fix: the cleanup path ran media_pattern.sub('') over the JSON-masked copy of the text, which baked the masking spaces into the user-visible 'cleaned' string — a serialized tool result like {"old":"MEDIA:/x.png"} came back as {"old":" "}. Now mask only a length-equal copy of 'cleaned' to locate the real tag spans, then delete those spans from the unmasked 'cleaned'. Real tags are stripped; JSON-embedded MEDIA: text reads back verbatim. Masking 'cleaned' (not the original 'content') keeps offsets valid after the [[audio_as_voice]] / [[as_document]] directives are removed. Adds two cleaned-text regression tests.	2026-05-31 23:51:42 -07:00
liuhao1024	e8827ef704	fix(gateway): skip MEDIA: inside serialized JSON string values Serialized tool results frequently embed a prior reply's text, e.g. {"result": "MEDIA:/path/stale.png"}. The bare-path branch of MEDIA_TAG_CLEANUP_RE matched these and re-delivered stale files (#34375). Adds BasePlatformAdapter._mask_json_string_media, which blanks (offset- preserving) only MEDIA:<bare-path> tokens that sit inside a JSON value- context string (opened by : , { or [). Legitimate tags at line start, after prose, indented, MEDIA:"quoted" form, and two-line TTS output are all left untouched. Reworked from the approach in #34388 (a line-start regex anchor), which no longer applied to current main and regressed same-line/indented tags. Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>	2026-05-31 23:51:42 -07:00
Nicolay	b3aaf2676b	fix(docker): discover Playwright headless_shell browser (#35717 ) Co-authored-by: Nic <nicsequenzy@gmail.com>	2026-06-01 16:06:44 +10:00
Ben Barclay	e3998d4714	chore(attribution): map polnikale for PR #35717 (#36273 ) Adds nicsequenzy@gmail.com -> polnikale to AUTHOR_MAP so the check-attribution gate passes for the Playwright headless_shell browser discovery fix (#35717).	2026-06-01 16:05:06 +10:00

1 2 3 4 5 ...

10214 Commits