PluginProbe ʕ •ᴥ•ʔ
AI Engine – The Chatbot, AI Framework & MCP for WordPress / 3.6.4
AI Engine – The Chatbot, AI Framework & MCP for WordPress v3.6.4
3.6.6 3.6.4 3.6.5 3.6.3 3.6.2 3.6.1 3.6.0 3.5.9 3.5.8 3.5.7 3.5.6 3.5.5 3.5.4 3.5.3 3.5.2 3.5.1 3.5.0 3.4.9 3.4.8 3.4.7 0.2.1 1.6.91 0.2.2 1.6.92 0.2.3 1.6.93 0.2.4 1.6.94 0.2.5 1.6.95 0.2.6 1.6.96 0.2.7 1.6.97 0.2.8 1.6.98 0.2.9 1.6.99 0.3.0 1.7.0 0.3.1 1.7.1 0.3.2 1.7.2 0.3.3 1.7.3 0.3.4 1.7.4 0.3.5 1.7.5 0.3.6 1.7.6 0.4.0 1.7.7 0.4.1 1.7.8 0.4.2 1.7.9 0.4.3 1.8.0 0.4.4 1.8.1 0.4.5 1.8.2 0.4.6 1.8.3 0.4.7 1.8.4 0.4.8 1.8.5 0.4.9 1.8.6 0.5.0 1.8.7 0.5.1 1.8.8 0.5.2 1.8.9 0.5.3 1.9.0 0.5.4 1.9.1 0.5.5 1.9.2 0.5.6 1.9.3 0.5.7 1.9.4 0.5.8 1.9.5 0.5.9 1.9.6 0.6.0 1.9.7 0.6.1 1.9.8 0.6.2 1.9.81 0.6.3 1.9.82 0.6.4 1.9.83 0.6.5 1.9.84 0.6.6 1.9.85 0.6.7 1.9.86 0.6.8 1.9.87 0.6.9 1.9.88 0.7.0 1.9.89 0.7.1 1.9.90 0.7.2 1.9.91 0.7.3 1.9.92 0.7.4 1.9.93 0.7.5 1.9.94 0.7.6 1.9.95 0.7.7 1.9.96 0.7.8 1.9.97 0.7.9 1.9.98 0.8.0 1.9.99 0.8.1 2.0.0 0.8.2 2.0.1 0.8.3 2.0.2 0.8.4 2.0.3 0.8.5 2.0.4 0.8.6 2.0.5 0.8.7 2.0.6 0.8.8 2.0.7 0.8.9 2.0.8 0.9.0 2.0.9 0.9.2 2.1.0 0.9.3 2.1.1 0.9.4 2.1.2 0.9.5 2.1.3 0.9.6 2.1.4 0.9.7 2.1.5 0.9.8 2.1.6 0.9.81 2.1.7 0.9.82 2.1.8 0.9.83 2.1.9 0.9.84 2.2.0 0.9.85 2.2.1 0.9.86 2.2.2 0.9.87 2.2.3 0.9.88 2.2.4 0.9.89 2.2.5 0.9.9 2.2.51 0.9.91 2.2.52 0.9.92 2.2.53 0.9.93 2.2.54 0.9.94 2.2.56 0.9.95 2.2.57 0.9.96 2.2.6 0.9.97 2.2.60 0.9.98 2.2.61 0.9.99 2.2.62 1.0.0 2.2.63 1.0.01 2.2.70 1.0.1 2.2.80 1.0.2 2.2.81 1.0.3 2.2.90 1.0.4 2.2.91 1.0.5 2.2.92 1.0.6 2.2.93 1.0.7 2.2.94 1.0.8 2.2.95 1.0.9 2.3.0 1.1.0 2.3.1 1.1.1 2.3.2 1.1.2 2.3.3 1.1.3 2.3.4 1.1.4 2.3.5 1.1.5 2.3.6 1.1.6 2.3.7 1.1.7 2.3.8 1.1.8 2.3.9 1.1.9 2.4.0 1.2.0 2.4.1 1.2.1 2.4.2 1.2.2 2.4.3 1.2.21 2.4.4 1.2.3 2.4.5 1.2.30 2.4.6 1.3.0 2.4.7 1.3.1 2.4.8 1.3.2 2.4.9 1.3.3 2.5.0 1.3.31 2.5.1 1.3.32 2.5.2 1.3.33 2.5.3 1.3.34 2.5.4 1.3.35 2.5.5 1.3.36 2.5.6 1.3.37 2.5.7 1.3.38 2.5.8 1.3.39 2.5.9 1.3.40 2.6.0 1.3.41 2.6.1 1.3.42 2.6.2 1.3.43 2.6.3 1.3.44 2.6.5 1.3.45 2.6.6 1.3.46 2.6.7 1.3.47 2.6.8 1.3.48 2.6.9 1.3.49 2.7.0 1.3.50 2.7.1 1.3.51 2.7.2 1.3.52 2.7.3 1.3.53 2.7.4 1.3.54 2.7.5 1.3.56 2.7.6 1.3.57 2.7.7 1.3.58 2.7.8 1.3.59 2.7.9 1.3.60 2.8.0 1.3.61 2.8.1 1.3.62 2.8.2 1.3.63 2.8.3 1.3.64 2.8.4 1.3.65 2.8.5 1.3.66 2.8.6 1.3.67 2.8.7 1.3.68 2.8.8 1.3.69 2.8.9 1.3.70 2.9.0 1.3.71 2.9.1 1.3.72 2.9.2 1.3.73 2.9.3 1.3.74 2.9.4 1.3.75 2.9.5 1.3.76 2.9.6 1.3.77 2.9.7 1.3.78 2.9.8 1.3.79 2.9.9 1.3.80 3.0.0 1.3.81 3.0.1 1.3.82 3.0.2 1.3.83 3.0.3 1.3.84 3.0.4 1.3.85 3.0.5 1.3.86 3.0.6 1.3.87 3.0.7 1.3.88 3.0.8 1.3.89 3.0.9 1.3.90 3.1.0 1.3.91 3.1.1 1.3.92 3.1.2 1.3.93 3.1.3 1.3.94 3.1.4 1.3.95 3.1.5 1.3.96 3.1.6 1.3.97 3.1.7 1.3.98 3.1.8 1.3.99 3.1.9 1.4.0 3.2.0 1.4.1 3.2.1 1.4.2 3.2.2 1.4.3 3.2.3 1.4.4 3.2.4 1.4.5 3.2.5 1.4.6 3.2.6 1.4.7 3.2.7 1.4.8 3.2.8 1.4.9 3.2.9 1.5.0 3.3.0 1.5.1 3.3.1 1.5.2 3.3.2 1.5.3 3.3.3 1.5.4 3.3.4 1.5.5 3.3.5 1.5.6 3.3.6 1.5.7 3.3.7 1.5.8 3.3.8 1.5.9 3.3.9 1.6.0 3.4.0 1.6.1 3.4.1 1.6.2 3.4.2 1.6.3 3.4.3 1.6.5 3.4.4 1.6.51 3.4.5 1.6.52 3.4.6 1.6.53 1.6.54 1.6.55 1.6.56 1.6.57 1.6.58 1.6.59 1.6.60 1.6.61 1.6.62 1.6.63 1.6.64 1.6.65 1.6.66 1.6.67 1.6.68 trunk 1.6.69 0.0.1 1.6.70 0.0.2 1.6.71 0.0.3 1.6.72 0.0.4 1.6.73 0.0.5 1.6.74 0.0.6 1.6.75 0.0.7 1.6.76 0.0.8 1.6.77 0.0.9 1.6.78 0.1.0 1.6.79 0.1.1 1.6.81 0.1.2 1.6.82 0.1.3 1.6.83 0.1.4 1.6.84 0.1.5 1.6.85 0.1.6 1.6.86 0.1.7 1.6.87 0.1.8 1.6.88 0.1.9 1.6.89 0.2.0 1.6.90
ai-engine / notes / QA-LOOP.md
ai-engine / notes Last commit date
CHANGELOG-PROPOSAL.txt 2 days ago QA-LOOP.md 2 days ago STUDY-CONSOLIDATION.md 2 days ago STUDY-MCP.md 2 days ago STUDY-WORKSPACE.md 2 days ago STUDY-WPAI-LEARNINGS.md 2 days ago
QA-LOOP.md
110 lines
1 # QA-LOOP.md
2
3 A periodic, self-driving QA loop for AI Engine. Each run picks **one fresh corner** of the
4 plugin, tries hard to break it, and either confirms it is solid or proposes a small fix.
5 Run it from time to time, let it tick every 30 minutes, and read the summaries.
6
7 ## How to run
8
9 ```
10 /loop 30m Run one QA-LOOP round on AI Engine (see QA-LOOP.md): pick an area NOT in the
11 rounds log below, try a feature/model/environment/UI/code path, be creative, and try to
12 trick it. Prioritise real user-facing friction over security. Report the result and, if you
13 find something, propose the smallest fix. Append the round to the log at the bottom of QA-LOOP.md.
14 ```
15
16 Stop it anytime by asking to stop the loop (the cron is session-only and also auto-expires after 7 days).
17
18 ## Why this loop exists
19
20 The thing that quietly kills a plugin is not a dramatic bug. It is a user who tries something,
21 it does not work nicely, and they drop it **without ever saying a word**. This loop hunts for
22 that: the rough edge, the corrupted output, the confusing error, the "why is this off by one line."
23
24 **Security is deliberately not the focus.** Multiple security teams already watch AI Engine.
25 Over-hardening just adds complexity and new edge cases for real users. If a round wanders into
26 security, note it and move back to friction and correctness.
27
28 ## Principles
29
30 - **One area per round.** Different from everything in the rounds log. Breadth over depth.
31 - **Try to trick it.** Malformed input, truncated streams, weird Unicode, huge numbers, empty
32 values, unusual-but-real model output. Think like a confused user or a cheap model, not an attacker.
33 - **Confirming robustness is a win.** Not every round needs a fix. "This corner is solid" is a
34 valuable, honest result. Do not manufacture a fix to feel productive.
35 - **Smallest possible fix.** If a fix is complex, propose the simpler alternative instead. Match
36 the surrounding code; reuse patterns already in the repo (e.g. the `min(500, …)` cap convention).
37 - **Do not churn deliberate behaviour alone.** When a behaviour is intentional and the call is a
38 product judgement, present options and ask rather than deciding unilaterally.
39 - **Leave dated breadcrumbs.** For deprecated/dead paths, compat shims, or deferred decisions, add
40 a `// TODO: Re-evaluate after <today + 6 months>` (absolute date, never a version number).
41 - **Prove it.** Reproduce logic in Node/PHP CLI, or test against the live site, before claiming a
42 result. Read code to confirm a suspicion; do not report speculation as fact.
43
44 ## Guardrails (hard rules)
45
46 - Never touch `Version:` / `MWAI_VERSION` / readme `Stable tag:` / the readme `== Changelog ==`
47 (all owned by Nekofy).
48 - Never partial-POST `settings/update` — it replaces the entire options object and wipes envs/keys.
49 Always read-modify-write the full object.
50 - Never commit without explicit approval. Group related changes into one coherent commit; past-tense,
51 one-line messages ending with ".". No mention of the assistant.
52 - Exclude build bundles from diffs (`app/chatbot.js`, `app/index.js`, `app/vendor.js`,
53 `premium/forms.js`, `premium/library-search.js`, etc.). Revert orphan bundle rebuilds that have no
54 matching source change.
55 - No em-dashes anywhere.
56 - Format PHP with `pcf fix <file>` (AI-Engine-only tool). Check the bundle mtime before browser-
57 verifying a JS edit; `pnpm build` once if stale.
58
59 ## Areas to rotate through
60
61 Pick one that is under-covered. This list is a starting point, not a limit; invent new angles.
62
63 - Chatbot rendering (markdown, code blocks, tables, lists, links, RTL, multibyte, emoji, HTML docs)
64 - Function calling / feedback loop / MCP tools (malformed args, multi-call, depth, timeouts)
65 - Models and providers (OpenAI, Anthropic, Google, OpenRouter, Perplexity, custom OpenAI-compatible)
66 - Environments and model resolution (bare model ids, wrong env, missing pricing)
67 - Embeddings / Knowledge / vector DBs (Pinecone, Qdrant, Chroma) / Smart Search
68 - AI Forms (placeholders, multi-file upload, submission)
69 - Discussions / conversation memory / title generation / history limits
70 - Usage stats / cost calculation / guest and user limits / the limit message UX
71 - Streaming / abort / mid-stream errors / stop button
72 - Content-aware / placeholders / templates
73 - Admin UI (settings, Playground, chatbot builder, Insights)
74 - Files / uploads / transcription / image generation and editing
75 - REST endpoints / parameter validation / sanitisation boundaries
76 - i18n and translatable strings
77
78 ## Test environment
79
80 - Live site: `https://ai.nekod.net/` (wp-admin available; local via `/etc/hosts`).
81 - Guest nonce: `POST /wp-json/mwai/v1/start_session``restNonce`.
82 - Chat submit: `POST /wp-json/mwai-ui/v1/chats/submit`.
83 - Guest usage limits are enforced; to bypass for a test, drop a temporary mu-plugin using the
84 `mwai_stats_credits` filter gated on a custom request header, and remove it right after.
85 - Env ids: OpenAI `9nx9mjyd`, Anthropic `e7944erg`, Google `q6ve6g9k`, Internal Test `intern01`.
86 - Read-only MCP tools (`wp_get_post`, `wp_get_posts`, `mcp_ping`, …) are safe for probing; never use
87 write/destructive MCP tools against the live site during a probe.
88
89 ## Report format (per round)
90
91 1. **Area + trick:** what was tested and how you tried to break it.
92 2. **Result:** solid (clean bill of health) or issue found.
93 3. **If an issue:** severity, the smallest fix or a proposal, and whether it was applied or deferred.
94 4. **Log it:** append a one-line entry below.
95
96 ## Rounds log
97
98 Append newest at the bottom. Keep entries to one line so future rounds can scan and avoid repeats.
99
100 - 2026-07-25 Gemini function-call result format + retired-model error hint — fixed both.
101 - 2026-07-25 Multibyte / CJK / emoji end-to-end pipeline — solid (mb-aware length, exact round-trip).
102 - 2026-07-25 Bare-URL underscore mangling in chatbot — fixed (URL placeholder protection).
103 - 2026-07-25 Content-aware injecting non-viewable posts — fixed (publicly-viewable / read_post guard).
104 - 2026-07-26 Markdown tables / lists rendering — fixed stray `<br/>` before nested lists; `1)` lists noted.
105 - 2026-07-26 Placeholder engine `[N/A]` catch-all — removed (corrupted content-aware pages + templates).
106 - 2026-07-26 MCP list-tool limits (`wp_get_posts` etc.) — capped at 500 to avoid big-site timeouts.
107 - 2026-07-26 Parameter validation (temperature / reasoning / maxTokens) — solid; temp clamp-to-1 noted.
108 - 2026-07-26 Usage cost for unknown / custom models — solid (records tokens, price null, no crash).
109 - 2026-07-26 Function-call streaming argument robustness — solid; legacy `function_call` path TODO'd.
110