Search by

stezkoy / flarum-ai-openreply

Stezkoy

AI Open-Reply Extension for Flarum, powered by opencode.

Package info

github.com/Stezkoy/flarum-ai-openreply

Issues

Forum

Type:flarum-extension

pkg:composer/stezkoy/flarum-ai-openreply

Statistics

Installs: 22

Dependents: 0

Suggesters: 0

Stars: 0

3.0.1 2026-09-28 18:20 UTC

This package is auto-updated.

Last update: 2026-09-28 18:22:35 UTC


README

License Latest Stable Version Total Downloads

A Flarum extension.

Automatically replies to new discussions (or to every post made by the original poster) using an AI assistant. Replies are generated by opencode — a headless opencode v2 server (opencode serve) — and posted by a designated assistant user.

Each discussion gets its own persistent opencode session, so the assistant keeps full context of the conversation. The assistant's reply is inserted into the thread as a text post.

opencode v2 required. Starting with version 3.0.0 this extension works only with the opencode v2 server (opencode serve, API under /api/*). The opencode v1 server API is no longer supported. See the server setup below.

A fork of michaelbelgium/flarum-ai-autoreply, reworked to work with the opencode v2 server API.

Requirements

  • Flarum >= 2.0 and PHP 8.3
  • A running opencode v2 server instance reachable from the Flarum host

Installation

Install with composer:

composer require stezkoy/flarum-ai-openreply

Then run migrations:

php flarum migrate

Installing opencode v2 on Ubuntu

Install the opencode v2 CLI on the machine that runs Flarum:

curl -fsSL https://opencode.ai/v2/install | bash

The script installs the opencode binary and prints next steps. Alternatively, install it globally via npm so the binary lands on your PATH (convenient for a systemd service later):

sudo npm install -g @opencode/cli
which opencode   # -> /usr/bin/opencode

If you previously installed the v1 package (opencode-ai), remove it first — v1 and v2 are not installed side by side, and the v2 installer replaces the v1 binary.

Configure at least one AI provider for opencode (this is done once per machine, not per session):

opencode auth login

Setting up the opencode server

The AI provider(s) and API keys are configured inside opencode itself. This extension only talks HTTP to the opencode server.

Safe default: localhost only (recommended)

The server binds to 127.0.0.1 by default, so it is only reachable from the same machine. Start it with basic auth and a fixed port:

OPENCODE_SERVER_PASSWORD=your-strong-password opencode serve --hostname 127.0.0.1 --port 49374

opencode v2 serves on port 49374 by default — set the extension's opencode server URL to http://localhost:49374 and you're done. Nothing is exposed to the network; you don't need to touch the firewall.

Exposing it externally (not recommended)

Only do this if the opencode server must live on a different machine than Flarum. Always set a password:

OPENCODE_SERVER_PASSWORD=your-strong-password opencode serve \
  --hostname 0.0.0.0 --port 49374 \
  --cors https://forum.example.com

Notes:

  • With --hostname 0.0.0.0 the server listens on all interfaces — restrict access with a firewall (ufw allow from <flarum-ip> to any port 49374) and put it behind TLS (Nginx/Caddy reverse proxy) if possible.
  • --cors is only needed for browser-based clients. This extension calls the API from PHP, so CORS is not required for it to work.

opencode serve options

Flag Description Default
--port Port to listen on 49374
--hostname Hostname to listen on 127.0.0.1
--cors Extra browser origins to allow (repeatable) []
--service Register as a background system service false
--stdio Communicate over stdio instead of HTTP false

Authentication is controlled with environment variables:

  • OPENCODE_SERVER_PASSWORD — enables HTTP basic auth (required if the server is reachable by anyone else).
  • OPENCODE_SERVER_USERNAME — basic auth username, defaults to opencode (the username the extension sends).

Running as a systemd service

Create /etc/systemd/system/opencode.service:

[Unit]
Description=opencode headless server (AI Open-Reply)
After=network-online.target
Wants=network-online.target

[Service]
Type=simple
User=opencode
WorkingDirectory=/var/lib/opencode
Environment=OPENCODE_SERVER_PASSWORD=your-strong-password
ExecStart=/usr/bin/opencode serve --hostname 127.0.0.1 --port 49374
Restart=always
RestartSec=3

[Install]
WantedBy=multi-user.target

Adjust ExecStart if opencode lives elsewhere (which opencode), then:

sudo useradd -r -s /usr/sbin/nologin --home-dir /var/lib/opencode opencode
sudo mkdir -p /var/lib/opencode && sudo chown opencode:opencode /var/lib/opencode
sudo systemctl daemon-reload
sudo systemctl enable --now opencode
sudo systemctl status opencode

The User=opencode account stores its own provider credentials and session data. Add Environment=OPENCODE_SERVER_USERNAME=... if you changed the basic auth username.

Running with Docker Compose

The easiest way to run the opencode server on a separate machine (for example, a VPS) is Docker. Official images are published on GHCR as ghcr.io/anomalyco/opencode. Pin an explicit version tag — the plain latest tag currently points to an opencode v1 release, so always use a 2.x tag (for example :2.0.0).

  1. On the VPS, install Docker Engine with the Compose plugin, then create the project directory:
sudo mkdir -p /opt/opencode && cd /opt/opencode
  1. Create compose.yaml:
services:
  opencode:
    image: ghcr.io/anomalyco/opencode:2.0.0
    container_name: opencode
    restart: unless-stopped
    ports:
      - "49374:49374"
    environment:
      OPENCODE_SERVER_PASSWORD: "${OPENCODE_SERVER_PASSWORD:?set it in .env}"
      OPENCODE_SERVER_USERNAME: "${OPENCODE_SERVER_USERNAME:-opencode}"
    volumes:
      - opencode-config:/root/.config/opencode
      - opencode-data:/root/.local/share/opencode
      - opencode-workspace:/workspace
    working_dir: /workspace
    command: ["serve", "--hostname", "0.0.0.0", "--port", "49374"]

volumes:
  opencode-config:
  opencode-data:
  opencode-workspace:
  1. Create .env next to it with the basic auth password:
OPENCODE_SERVER_PASSWORD=your-strong-password
# OPENCODE_SERVER_USERNAME=opencode   # optional; the default matches the extension's username
  1. Start the server, check the logs, and log into your AI provider (interactive, once):
sudo docker compose up -d
sudo docker compose logs -f opencode   # confirm the server is up (Ctrl+C to exit)
sudo docker compose exec -it opencode opencode auth login

What this does:

  • The container runs opencode serve --hostname 0.0.0.0 --port 49374 — it listens on all interfaces on port 49374, with basic auth enabled via OPENCODE_SERVER_PASSWORD (username opencode).
  • The opencode-config and opencode-data volumes persist your provider credentials (auth.json) and all discussion sessions on the VPS. They survive docker compose down and image upgrades; only docker compose down -v deletes them — don't run that if you want to keep sessions.
  • opencode-workspace is mounted as /workspace (working_dir) — the project the server reports to the API.

Then point the extension's admin settings at the server:

  • opencode server URL — http://<vps-ip>:49374 (replace <vps-ip> with the VPS address)
  • opencode server username — opencode
  • opencode server password — the value of OPENCODE_SERVER_PASSWORD

Security:

  • Open the port only to the host running Flarum: sudo ufw allow from <flarum-ip> to any port 49374 proto tcp. Basic auth alone does not protect against brute-force password guessing, and the API can spend money generating replies.
  • Prefer TLS in front of it (Caddy/Nginx reverse proxy) if the VPS is reachable from the internet.
  • --cors is not needed — the extension calls the API from PHP, not from a browser.

Upgrading:

cd /opt/opencode && sudo docker compose pull && sudo docker compose up -d

To move to a newer release, bump the image tag in compose.yaml (for example :2.0.0 → :2.0.1) and run the command above. Sessions and credentials survive the upgrade because they live in the volumes.

Server recommendations

opencode is a local LLM agent runtime: it loads the model, keeps the full conversation (and often the context window) in memory, and may spawn long-running background tasks. The headless server used by this extension is no exception, so plan its resources accordingly.

Performance and stability

  • Run it next to Flarum on a dedicated machine (or container), not on the same process as PHP. Reply generation can be slow (tens of seconds to minutes), so isolate it from the web stack and give it its own CPU/memory budget and network.
  • Ensure the extension can reach the server. The extension uses a 600-second request timeout (OpencodeClient), so slow generations are fine, but the connection timeout is 5 seconds — the server must respond quickly at the TCP level. Keep them on the same LAN / low-latency network to avoid spurious failures.
  • Watch memory. Each discussion keeps its own persistent session; long conversations and large context windows accumulate in RAM. Set a safe limit on the number of active discussions or restart the service periodically (Restart=always above) to reclaim memory.
  • Use a fast network and a stable provider for the upstream LLM. generation latency and provider rate limits dominate response time.
  • Apply an access token / basic auth so the server is not left open (see authentication section above).

Minimum / recommended hardware

These are rough guidelines for a CPU-only setup running a headless server that serves a small-to-medium forum. Rent a dedicated VM or container rather than sharing a tiny VPS with heavy applications.

Model size (approx.) RAM vCPU Notes
Small (cloud / hosted APIs) 2–4 GB 1–2 Model runs on the provider, little local load
Local 7–8B quantized model 16 GB 4 Comfortable for multi-session use
Large 70B+ / many concurrent 64 GB+ 8+ Heavy; prefer hosted APIs instead

If you are using hosted APIs (recommended for most forums), opencode itself only needs modest resources — the table's first row. Save the bigger rows for running local models.

In the extension's admin settings page:

  • opencode server URL — the address of your headless opencode server (default http://localhost:49374, matching the v2 server's default port).
  • opencode server username — the basic auth username (default opencode). Used only when a password is set.
  • opencode server password — the OPENCODE_SERVER_PASSWORD value if basic auth is enabled.
  • Agent — a preset that shapes how the assistant replies: the standard build (answers right away) or plan (thinks the answer over first), or the default. The server-side validation falls back to the default agent (with a server-log warning) if a saved name is unknown, so a reply never fails on a typo.
  • System prompt (persona) — optional free-text instructions for the assistant's behavior, e.g. "call yourself Pupsik and answer in Russian". Stored with each discussion's opencode session (the session's instructions, opencode 2.x) and applied as part of the model's context — so the persona stays out of the visible messages and can differ per discussion.
  • Model — the model to use, in provider/model format (e.g. opencode/big-pickle). Type it manually, or click Get free models to fetch the currently available free models from your opencode server and click one to fill the field. Leave empty to use the server's default model. This is independent of the agent: the agent fixes how it behaves, the model fixes which AI answers.
  • User assistant — the user ID of the account that posts the AI replies (required).
  • User assistant badge — a toggle plus the text shown in the badge below the assistant's posts. Disabled or empty — no badge is rendered.
  • Enable on discussion start — when enabled, the AI replies only when a discussion is started. When disabled, the discussion becomes a chat between the OP and the assistant.
  • Tags — restrict the assistant to specific tags.
  • Actions — three buttons: Check connection (server health + current model), Count sessions (total sessions on the server and how many belong to this extension), and Close all sessions (closes this extension's sessions in bounded batches — if some remain, click again; other server sessions are left untouched).
  • Resource limits — max_active_sessions, max_messages_per_session, session_ttl_days. Set 0 to disable a limit.
  • Retries — retry_attempts (total attempts, default 1, max 10) and retry_delay_seconds (delay before each retry, default 1, max 120) for requests to the opencode server.

The agent and model are applied when a session is created, so new discussions pick up the latest settings. The assistant's persona can be set directly in the extension's admin page above (it lands in the discussion session's instructions); alternatively the persona can be baked into the agent itself on the server, e.g. in opencode.json:

{
  "agents": {
    "build": {
      "system": "You are a helpful assistant on a Flarum forum. Answer in the language of the user's post."
    }
  }
}

Also grant the "Use AI assistant" permission to the desired user groups.

Features

  • Auto-reply to new discussions using AI
  • Chat mode: a discussion becomes a 1-on-1 chat between the OP and the assistant
  • Persistent per-discussion context (one opencode session per discussion)
  • Replies are posted as regular text posts by a designated assistant user
  • Restrict the assistant to selected tags
  • Permission controls for who can trigger the auto-reply
  • Replies go through the full Flarum event lifecycle: discussion counters, last-post pointers and subscriber notifications stay correct
  • Dynamic free model list fetched live from the opencode server (no hardcoded presets)
  • Configurable retries (retry_attempts, retry_delay_seconds) and resource limits

Updating

Important: version 3.x works only with the opencode v2 server. Before updating, make sure your opencode server is v2 (JSON API under /api/*, default port 49374). The opencode 1.x server API is no longer supported.

composer update stezkoy/flarum-ai-openreply
php flarum migrate
php flarum cache:clear

Links