Files
MyAiDesk/nomifun-tauri/docs/guides/channels.md
T
freedak f7a720204a Update: 将子项目从 submodule 转为完整内容
- 移除 GovAI, nomifun-tauri, 算力盒子 的 submodule 引用
- 添加所有子项目的完整源代码
- 保留原始 .git 为 .git.bak 备份
2026-07-04 19:20:46 +08:00

14 KiB

Channels

A channel lets you operate a NomiFun agent from an external chat app — Telegram, Lark / 飞书, DingTalk, WeChat — instead of sitting in front of the desktop window. You enable a plugin, paste in its credentials, authorize a chat user with a one-time code, and from then on messages to your bot are dispatched to the agent and its replies come back into the same thread.

Channels are useful when:

  • you want to brief an agent from your phone or a group chat;
  • you want a workspace-aware agent reachable from a team's existing IM;
  • you want long-running tasks (AutoWork) to be kickable from outside the desktop without spinning up the WebUI.

Each platform plugin is a Cargo feature on nomifun-channel (telegram, lark, dingtalk, weixin). The default NomiFun build ships with all of them on; if you build the backend yourself with a non-default feature set, the corresponding tab simply disappears.

Channels settings overview

Where to find it

Open the Nomi page (/nomi), select a companion, and switch to the Remote tab (/nomi?companion=<id>&tab=remote). That tab lists the remote connectors for the selected companion — built-in (Telegram, Lark, DingTalk, WeChat, WeCom, Slack, Discord, extensions). For each plugin you'll see:

  • a status pill (stopped / connected),
  • the bot username once connected,
  • the number of currently authorised users,
  • a per-channel default agent + default model selector.

Slack / Discord / WeCom appear as built-in placeholders today — the backend wiring is feature-gated and still being built out for those two; Telegram / Lark / DingTalk / WeChat are the ones you can run today.

How a channel works

external IM ──▶ plugin (long-poll / WebSocket)
                    │
                    ▼
            ChannelManager  ◀─▶  PairingService
                    │
                    ▼
              SessionManager  ──▶  agent / conversation
  • Plugin owns the platform-specific connection (Telegram long-poll with exponential backoff, Lark / DingTalk WebSocket, WeChat QR-code login over SSE).
  • PairingService turns "I'm John on Telegram, let me in" into a 6-digit code that you approve from the desktop UI.
  • SessionManager maps (platform_user, chat_id) to an agent conversation, so each external chat is a stable session and follow-up messages land in the same agent.
  • Orchestrator plumbs incoming messages into the agent stream and the agent's replies back out as edits to the same IM message (everything except WeChat supports message editing — WeChat falls back to sending follow-up replies).

Setting up each platform

Telegram

  1. Talk to @BotFather and create a bot. Save the token (looks like 123456:ABC-DEF…).
  2. In Nomi → Remote → Telegram, paste the token.
  3. Click Test — the backend calls getMe and shows the bot username on success.
  4. Click Enable. The plugin starts long-polling (25 s timeout, exponential backoff up to 10 reconnects).

To pair a Telegram user with the desktop, the user messages your bot; the bot replies with a 6-digit code (10-minute TTL). Paste / type the code into Nomi → Remote → Pending pairings on the desktop and click Approve. From then on that Telegram user can chat with the agent.

Lark / Feishu

  1. Create a custom app in the Lark developer console with the events you need (text message, card action, bot menu).
  2. Copy the App ID, App Secret, and (optional) Encrypt key / Verification token.
  3. Paste them into the Lark form in the Channels tab and click Enable.

The Lark plugin connects via Lark's WebSocket long-connection (no public webhook needed), with a 60-second event-dedup cleanup loop and fragment reassembly. Replies are sent as interactive cards because Lark's API only supports editing card messages.

DingTalk

  1. Create an internal app in DingTalk Developer Backstage with Stream Mode enabled.
  2. Copy the Client ID and Client Secret into the DingTalk form and enable.

The DingTalk plugin opens a WebSocket using the standard DingTalk stream-mode handshake; pairing flow is identical to Telegram.

WeChat

  1. WeChat is QR-code login. Click Enable on the WeChat plugin — the backend opens an SSE stream (POST /api/channel/weixin/login/start) that pushes QR-code refresh events.
  2. Scan the QR with the WeChat app, confirm the login, and the plugin transitions to connected.

WeChat does not support message editing — replies are delivered as new messages in the same chat instead of in-place edits.

Pairing and authorising users

A pairing request comes in two ways:

  1. The platform user messages the bot for the first time (Telegram /Lark / DingTalk). The plugin auto-creates a pending request and replies to the user with the code.
  2. You can approve / reject the pending request from Nomi → Remote → Pending pairings or programmatically via POST /api/channel/pairings/approve and POST /api/channel/pairings/reject.

Approved users are listed in Authorised users, with last active. You can revoke at any time (POST /api/channel/users/revoke); the service also cleans up that user's open sessions so the next message re-pairs from scratch.

Pairing approval

Master Agent mode

By default, every channel conversation runs in Master Agent mode: the remote message is greeted by the Nomi companion itself. The conversation inherits the companion's personality and memories, and the agent is wired to the Desktop Gateway tools, so from your phone you're not talking to an isolated chat bot — you're talking to the agent that runs your desktop.

What the gateway tools (all prefixed nomi_*, 32 of them today) let the remote agent do on your behalf:

  • Conversations — list every conversation with its runtime state, inspect one (status plus the latest messages, including an in-flight streaming reply), send a message or task prompt into any conversation, create new ones, update or delete old ones (nomi_list_conversations, nomi_conversation_status, nomi_send_to_conversation, nomi_create_conversation, nomi_update_conversation, nomi_delete_conversation).
  • Scheduled tasks — list / create / update / delete cron jobs (nomi_cron_list, nomi_cron_create, nomi_cron_update, nomi_cron_delete).
  • Long-term memory — read and write the companion's global memory bank (nomi_memory_list, nomi_memory_save, nomi_memory_update, nomi_memory_delete).
  • Requirements — browse and manage the requirements platform (nomi_requirement_list, nomi_requirement_create, nomi_requirement_update, nomi_requirement_delete).
  • Terminals & supervision — list terminal sessions, create new ones (optionally binding knowledge bases via knowledge_base_ids), and read / toggle a terminal's AutoWork binding and IDMM supervision (nomi_list_terminals, nomi_create_terminal, nomi_get_autowork, nomi_set_autowork, nomi_get_idmm, nomi_set_idmm).
  • Knowledge bases — browse bases and bindings, rebind a conversation / terminal / companion, create a new base, write markdown files into one, trigger the AI digest, or fetch a URL as markdown — so the companion can deposit knowledge on its own (nomi_knowledge_list_bases, nomi_knowledge_get_binding, nomi_knowledge_set_binding, nomi_knowledge_create_base, nomi_knowledge_write_file, nomi_knowledge_autogen, nomi_knowledge_fetch_url). nomi_knowledge_create_base with urls fetches in the background — the call returns immediately, so don't create the base a second time while waiting; the base's description appearing means the fetch + digest pipeline is done.
  • Providers — list the configured LLM providers (nomi_list_providers).

So "move my daily-report cron to 9 am and tell me what's running right now" is a single Lark message.

Turning it off. Each platform panel has a Master Agent mode switch next to the default-model selector. It's on by default; the preference is stored per platform as assistant.<platform>.masterAgent in the client preferences (missing value = on). Switching it off reverts that platform to the legacy behavior — each remote chat gets a plain standalone conversation, with no companion persona and no gateway tools. Like the model selector, toggling the switch calls POST /api/channel/settings/sync and clears the platform's active sessions, so the next inbound message starts a conversation in the new mode.

Choosing which companion greets the channel. With multiple companions, bots are bound to companions per channel row: each row of assistant_plugins is one bot (the same platform can host several — e.g. one Feishu in-house app per companion), its companion_id decides which companion answers, and the UNIQUE(type, bot_key) constraint structurally guarantees one bot is never bound to two companions (bot identity: Feishu app_id, the Telegram bot id, DingTalk client_id, …). Binding or unbinding calls POST /api/channel/settings/companion with a plugin_id, which persists the row and resets that channel's active sessions in one step — the next inbound message is greeted by the new companion's persona, model, and knowledge mounts (the conversation carries extra.companionId). Connecting a bot from a companion's Remote tab creates the channel row and binds it to that companion in one go. A row without a companion binding falls back to the legacy per-platform preference assistant.<platform>.companionId, then to the default companion; if the bound companion is later deleted, the channel falls back to the default companion and the sessions are likewise reset. Memory is shared across the whole companion family: no matter how many bots and channels you connect, their conversations flow into the same single memory pipeline, so switching companions never loses memories.

How it relates to the agent / model pickers. The per-platform Default agent still decides which engine answers; the gateway tools are injected for any agent type, while the companion persona and memory ride on the Nomi engine. Model resolution in master mode: the platform's Default model (if set) wins, otherwise the conversation falls back to the bound companion's own model.

Picking the agent and model

Each platform has a Default agent and Default model selector in its config form. The platform stores them as assistant.<platform>.defaultModel in the client config, so:

  • a message from Telegram routes to whatever agent / model you picked for Telegram;
  • a message from Lark can route to a different agent;
  • changing the selector calls POST /api/channel/settings/sync, which clears any active sessions for that platform — the next inbound message re-creates them with the new defaults.

The model selector is the same Gemini-flavoured component the desktop uses, so any provider you've configured (Anthropic, OpenAI-compatible custom URL, Gemini-with-Google-auth, Bedrock, …) is available here.

What works from the IM side

The platform-agnostic abstraction (UnifiedIncomingMessage, UnifiedOutgoingMessage, UnifiedAction) covers:

  • Plain text — both directions.
  • Edited streaming responses — incremental updates from the agent are edited into the in-flight bot message (not on WeChat).
  • Action buttons — confirmation prompts, retry actions, etc., rendered as inline keyboards (Telegram), interactive-card buttons (Lark), or platform equivalents.
  • Bot mention / require-mention — group chats can be configured to only respond when the bot is @-mentioned.

What you don't get from the IM side (yet):

  • spawning teams (use the desktop / web UI for that);
  • file uploads beyond what the platform plugin natively understands;
  • per-user workspace selection — the agent's workspace is the one set on the conversation it routed to.

Routes & API

What Where
Channels UI /nomi?companion=<id>&tab=remote
List plugins / status GET /api/channel/plugins
Enable / disable POST /api/channel/plugins/enable, …/disable
Test credentials POST /api/channel/plugins/test
Pending pairings GET /api/channel/pairings
Approve / reject pairing POST /api/channel/pairings/approve, …/reject
Authorised users GET /api/channel/users, POST .../users/revoke
Active sessions GET /api/channel/sessions
Sync (clear sessions on change) POST /api/channel/settings/sync
Bind master-agent companion POST /api/channel/settings/companion
WeChat QR login SSE POST /api/channel/weixin/login/start

Notes

  • Plugin lifecycle is a state machine — Created → Initializing → Ready → Starting → Running → Stopping → Stopped, with any step able to transition to Error. The status pill in the UI is this enum.
  • A revoked user's session is torn down before the user row is deleted. The next message from that platform user will trigger a new pairing code.
  • Pairing codes are 6 digits, generated with getrandom, with a 10-minute TTL. The pairing service runs a periodic sweep that expires pending codes whose TTL has passed.
  • WeChat is feature-gated separately because its dependency tree is heavier (QR / login / auth flow). If you build with --no-default-features, you'll see the placeholder card but no enable button.
  • Companions — multi-companion management, shared memory, and the per-companion knowledge bindings that ride on channel conversations.
  • AutoWork & Requirements — file a requirement from a chat, get notified when it lands via a webhook to Lark / HTTP / Slack (configured at 需求平台 → 扩展能力 → 通知).
  • Web Server Deployment — exposes the same channels when you self-host the backend on a server.