PluginProbe
ThinkRank AI SEO – AI SEO Plugin for WordPress: Schema, XML Sitemaps, Meta Tags, Search Console & Local SEO / 2.14.1
ThinkRank AI SEO – AI SEO Plugin for WordPress: Schema, XML Sitemaps, Meta Tags, Search Console & Local SEO v2.14.1
2.14.2 2.14.1 2.14.0 2.13.0 2.12.0 2.11.0 2.10.0 2.9.0 2.8.0 2.7.0 2.6.0 2.5.0 2.4.0 2.3.0 2.2.0 2.1.1 2.1.0 2.0.2 2.0.1 2.0.0 1.32.0 1.31.0 1.30.0 1.29.0 1.28.0 All 57 releases
thinkrank / includes / seo / class-builder-content.php

class-builder-content.php in ThinkRank AI SEO – AI SEO Plugin for WordPress: Schema, XML Sitemaps, Meta Tags, Search Console & Local SEO 2.14.1, at includes/seo/class-builder-content.php

2,136 lines 82.8 KB
No matching file
Up and down to move Enter to open Esc to close
Raw Download Zip
1 <?php
2 /**
3 * Page-builder content extraction.
4 *
5 * SEO analysis reads `post_content`, which is only the real text on a classic
6 * post. Page builders keep the words somewhere else, and every server-side
7 * scoring path — bulk analysis, the post-list SEO Overview column, the MCP
8 * abilities, cron reports — saw an empty page as a result:
9 *
10 * - Oxygen / Breakdance leave `post_content` completely EMPTY and store the
11 * node tree in postmeta. Nothing to render, nothing to strip: the analyzer
12 * reported "No content" on pages with well over a thousand visible words.
13 * - Elementor does the same via `_elementor_data`.
14 * - Divi 5 and Gutenberg do store block markup in `post_content`, but Divi
15 * keeps module text inside the block's JSON attributes — inside an HTML
16 * comment, which tag stripping removes wholesale.
17 * - Divi 4 and other shortcode builders keep text in shortcode attributes.
18 *
19 * Extraction reads the builder's own stored data rather than invoking its
20 * render engine. Rendering an Oxygen page outside a front-end request is slow,
21 * stateful and can fatal in an admin context, whereas the stored tree is just
22 * JSON — cheap, side-effect free and safe to touch during a bulk run.
23 *
24 * @package ThinkRank\SEO
25 * @since 1.23.0
26 */
27
28 declare(strict_types=1);
29
30 namespace ThinkRank\SEO;
31
32 if (!defined('ABSPATH')) {
33 exit;
34 }
35
36 /**
37 * Resolves the analyzable content of a post, whatever built it.
38 */
39 class Builder_Content {
40
41 /**
42 * Post meta keys that hold builder data, in priority order.
43 *
44 * Several generations of the same builder are listed on purpose. Oxygen 6
45 * is Breakdance under the hood and writes the same tree, under its own
46 * prefix: Breakdance keeps it in `_breakdance_data`, Oxygen 6 in
47 * `_oxygen_data` (the key is `__bdox('_meta_prefix') . 'data'`, and the
48 * prefix is `_oxygen_` under Oxygen). Both store it inside a
49 * `tree_json_string` envelope, see unwrap_tree_envelope(). Earlier Oxygen
50 * releases used the shortcode-based `ct_builder_shortcodes` and its JSON
51 * sibling. A site can only have one of them.
52 *
53 * @var string[]
54 */
55 private const BUILDER_META_KEYS = [
56 '_breakdance_data', // Breakdance
57 '_oxygen_data', // Oxygen 6+ (Breakdance engine, Oxygen prefix)
58 // Oxygen classic. 4.x writes the tree as JSON to `ct_builder_json`
59 // while still keeping `ct_builder_shortcodes`. A post carrying only
60 // the JSON key used to match no key at all and fall through to an
61 // empty `post_content`, which reads as a one-word page (#776).
62 //
63 // Oxygen 4.8.3 then renamed every `ct_*` post meta key to `_ct_*`
64 // (`oxygen_vsb_update_4_8_3()` runs `oxy_prefix_meta_keys()` on
65 // upgrade, and `oxy_get_post_meta()` only reads the prefixed name
66 // from then on). A current Oxygen classic site therefore has only the
67 // underscored keys, which nothing here listed, so every one of its
68 // pages resolved as empty. The prefixed keys come first because they
69 // are what Oxygen itself reads; the bare ones cover a site that has
70 // not run the migration (or reverted it with `?unprefix_meta`).
71 //
72 // These four are not read in this loop: from_oxygen_classic() pairs
73 // each JSON key with its shortcode sibling so the two forms can be
74 // compared. They are listed here because this list is also what the
75 // word-count index watches and what FAQ detection scans.
76 '_ct_builder_json', // Oxygen classic 4.8.3+ (JSON tree)
77 'ct_builder_json', // Oxygen classic 4.0-4.8.2 (JSON tree)
78 '_ct_builder_shortcodes', // Oxygen classic 4.8.3+ (shortcode tree)
79 'ct_builder_shortcodes', // Oxygen classic < 4.8.3 (shortcode tree)
80 '_elementor_data', // Elementor
81 // Beaver Builder. Published layout first: `_fl_builder_draft` holds
82 // unsaved changes and would score content the visitor cannot see.
83 // Both are arrays of stdClass nodes, which is why the walker below
84 // has to treat objects like arrays (#449).
85 '_fl_builder_data', // Beaver Builder (published)
86 '_fl_builder_draft', // Beaver Builder (unsaved changes)
87 ];
88
89 /**
90 * Bricks' content-area meta key, used when Bricks itself isn't loaded.
91 *
92 * Bricks exposes `BRICKS_DB_PAGE_CONTENT` and renames the underlying key
93 * between generations (it gained the `_2` suffix in 1.7.3), so the
94 * constant is authoritative and this literal is only the fallback for the
95 * contexts where it is undefined — Bricks is a theme, so on an admin or
96 * CLI request against a site that has since switched themes the constant
97 * is simply not there while the post meta still is.
98 *
99 * Bricks stores three areas — header, content and footer. Only the content
100 * area belongs to the post being scored; the header and footer areas live
101 * on Bricks' own template posts and would double-count site chrome into
102 * every page's word count, so they are deliberately not read here.
103 *
104 * @since 2.2.1
105 * @var string
106 */
107 private const BRICKS_CONTENT_META_KEY = '_bricks_page_content_2';
108
109 /**
110 * Bricks' per-post editor-mode meta key, used when Bricks isn't loaded.
111 *
112 * @since 2.2.1
113 * @var string
114 */
115 private const BRICKS_EDITOR_MODE_META_KEY = '_bricks_editor_mode';
116
117 /**
118 * Bricks' components option, used when Bricks itself isn't loaded.
119 *
120 * @since 2.2.1
121 * @var string
122 */
123 private const BRICKS_COMPONENTS_OPTION = 'bricks_components';
124
125 /**
126 * Bricks' element that renders the post's own `post_content`.
127 *
128 * A Bricks page normally discards `post_content` entirely, which is why
129 * anything left there is invisible. Dropping this element onto the canvas
130 * is the one way an author puts it back on the page, so its presence flips
131 * `post_content` from stale leftovers to content the visitor reads.
132 *
133 * @since 2.3.1
134 * @var string
135 */
136 private const BRICKS_POST_CONTENT_ELEMENT = 'post-content';
137
138 /**
139 * Bricks' Heading element, and the tag it renders when none is stored.
140 *
141 * Bricks leaves a setting out of storage while it equals its default, so a
142 * Heading left on its default tag is stored with no `tag` at all. Bricks
143 * 2.4.1 renders it as `h3` (`Element_Heading::$tag`, overridable by the
144 * active theme style's `tag`), and the walker, which only wraps text whose
145 * node names a tag, read it as body copy (#908).
146 *
147 * @since 2.15.0
148 * @var string
149 */
150 private const BRICKS_HEADING_ELEMENT = 'heading';
151
152 /**
153 * Tag a Bricks Heading renders when neither it nor a theme style sets one.
154 *
155 * @since 2.15.0
156 * @var string
157 */
158 private const BRICKS_HEADING_DEFAULT_TAG = 'h3';
159
160 /**
161 * Tag an Elementor Heading widget renders when `header_size` is not stored.
162 *
163 * Elementor saves `settings.toJSON({ remove: ['default'] })`, so a heading
164 * left on its default size has no `header_size` in `_elementor_data`, and
165 * that default is `h2` (#908).
166 *
167 * @since 2.15.0
168 * @var string
169 */
170 private const ELEMENTOR_HEADING_DEFAULT_TAG = 'h2';
171
172 /**
173 * Resolved Bricks trees for this request, keyed by post ID.
174 *
175 * Rendering one page asks for the tree about twenty times — every
176 * description, every schema node, the FAQ guard — and resolving it is not
177 * free. `bricks_content_source()` clears `Bricks\Database::$active_templates`
178 * before asking Bricks which content template applies, which defeats
179 * Bricks' own early-return and re-runs its whole template-condition engine;
180 * `expand_bricks_components()` then walks the tree again. Measured on a
181 * Bricks page with no content of its own, that was ten full runs of the
182 * rules engine per request.
183 *
184 * Per-request only, and only ever read back within one page render — a
185 * request that writes Bricks content does not also render it.
186 *
187 * @since 2.3.1
188 * @var array<int,array<int,mixed>>
189 */
190 private static array $bricks_trees = [];
191
192 /**
193 * JSON keys whose values are user-visible text.
194 *
195 * Builder trees mix content with configuration, so a blind string sweep
196 * would count CSS classes and option slugs as words. Matching on the key
197 * keeps the word count honest.
198 *
199 * @var string[]
200 */
201 private const CONTENT_KEYS = [
202 'text', 'title', 'subtitle', 'heading', 'subheading', 'content',
203 'description', 'caption', 'excerpt', 'label', 'value', 'html',
204 'editor', 'quote', 'answer', 'question', 'body', 'button_text',
205 // Oxygen classic keeps an element's copy in `options.ct_content`
206 // (headline, text block, rich text, link and button labels). It is the
207 // field Oxygen's own serializer moves between the tags when it writes
208 // shortcodes (`parse_components_tree()`), and the one Relevanssi and
209 // Oxygen's WPML integration read. Missing from this list, the walker
210 // kept only copy that happened to contain markup: a page of plain
211 // headings and paragraphs lost almost all of its words.
212 'ct_content',
213 // Oxygen's composite elements keep their copy under `options.original`
214 // instead, one key per field. Taken from the list Oxygen itself treats
215 // as text when it serializes (`$options_to_encode`); the numeric price
216 // fields and the progress bar's right-hand percentage are left out.
217 'testimonial_text', 'testimonial_author', 'testimonial_author_info',
218 'icon_box_heading', 'icon_box_text',
219 'pricing_box_package_title', 'pricing_box_package_subtitle', 'pricing_box_content',
220 'progress_bar_left_text',
221 ];
222
223 /**
224 * Oxygen classic's storage generations, as JSON key => shortcode key.
225 *
226 * @since 2.10.0
227 * @var array<string,string>
228 */
229 private const OXYGEN_CLASSIC_KEYS = [
230 '_ct_builder_json' => '_ct_builder_shortcodes',
231 'ct_builder_json' => 'ct_builder_shortcodes',
232 ];
233
234 /**
235 * JSON keys whose values hold a link destination.
236 *
237 * Builders store a link's destination in a structured field separate from
238 * its label, either as a bare URL string or as a `{ url: … }` object.
239 * Neither shape survives a text sweep — the key is not content and a bare
240 * URL contains no `<` — so no `<a>` tag reached the link counters.
241 *
242 * @var string[]
243 */
244 private const URL_KEYS = [
245 'link', 'url', 'href', 'link_url', 'button_link', 'permalink', 'link_to',
246 ];
247
248 /**
249 * JSON keys whose values hold an embedded video's source.
250 *
251 * A builder's video widget keeps its destination in a provider-specific
252 * field — Elementor picks `youtube_url`, `vimeo_url`, `dailymotion_url` or
253 * `hosted_url` according to the chosen source type — none of which is a
254 * link field or a content field, so a video on a builder page reached the
255 * analyzers as nothing at all.
256 *
257 * These are deliberately kept out of URL_KEYS. A video is an embed, not an
258 * outbound link: rendering one as `<a href>` would add a spurious external
259 * link to every page carrying a video and skew the link counts. They are
260 * reconstructed as `<iframe>`/`<video>` instead, which the video detector
261 * recognises and the link and image counters ignore.
262 *
263 * @since 2.3.1
264 * @var string[]
265 */
266 private const VIDEO_KEYS = [
267 'youtube_url', 'vimeo_url', 'dailymotion_url', 'videopress_url',
268 'hosted_url', 'video_url', 'video_src', 'video_link',
269 ];
270
271 /**
272 * File extensions that mean a video source is a file, not a provider page.
273 *
274 * @since 2.3.1
275 * @var string[]
276 */
277 private const VIDEO_FILE_EXTENSIONS = ['mp4', 'webm', 'ogv', 'mov', 'm4v'];
278
279 /**
280 * Source keys to trust for a declared video source type.
281 *
282 * A widget keeps one field per provider and does not clear the others when
283 * the author switches source: an Elementor video moved from YouTube to Self
284 * Hosted still carries the earlier `youtube_url`. Reading whichever key
285 * turns up first then emits the video the author replaced. The widget says
286 * which one it is actually playing, so that is read first and the flat key
287 * sweep is only the fallback for a builder that declares nothing.
288 *
289 * @since 2.3.1
290 * @var array<string,string[]>
291 */
292 private const VIDEO_KEYS_BY_TYPE = [
293 'youtube' => ['youtube_url'],
294 'vimeo' => ['vimeo_url'],
295 'dailymotion' => ['dailymotion_url'],
296 'videopress' => ['videopress_url'],
297 'hosted' => ['hosted_url', 'video_url', 'video_src', 'video_link'],
298 'media' => ['hosted_url', 'video_url', 'video_src', 'video_link'],
299 'file' => ['hosted_url', 'video_url', 'video_src', 'video_link'],
300 'self_hosted' => ['hosted_url', 'video_url', 'video_src', 'video_link'],
301 ];
302
303 /**
304 * Keys a builder uses to name which video source a widget is playing.
305 *
306 * @since 2.3.1
307 * @var string[]
308 */
309 private const VIDEO_TYPE_KEYS = ['video_type', 'videotype', 'video_source', 'source_type'];
310
311 /**
312 * JSON keys whose values hold an image, as a URL string or `{ url, alt }`.
313 *
314 * @var string[]
315 */
316 private const IMAGE_KEYS = [
317 'image', 'src', 'image_url', 'background_image', 'bg_image', 'photo',
318 ];
319
320 /**
321 * JSON keys that carry a heading level for the node's text.
322 *
323 * A builder heading's text is collected (its key is in CONTENT_KEYS) and so
324 * counts toward the word count, but it arrives as bare text with no `<h2>`
325 * wrapper — which is why heading-structure checks saw none.
326 *
327 * @var string[]
328 */
329 private const HEADING_TAG_KEYS = [
330 // Lower-cased on both sides of the comparison, so `headingtag` is the
331 // camelCase `headingTag` Bricks uses throughout its own controls and
332 // which ThinkRank's Bricks elements declare. Without it their section
333 // headings counted as body copy and never reached heading structure.
334 'header_size', 'heading_tag', 'headingtag', 'html_tag', 'title_tag', 'tag', 'level', 'size',
335 ];
336
337 /**
338 * Keys whose value is alternative text for a sibling image.
339 *
340 * @var string[]
341 */
342 private const ALT_KEYS = ['alt', 'alt_text', 'image_alt', 'title'];
343
344 /**
345 * The global post and every global `setup_postdata()` writes.
346 *
347 * Rendering points them at the post being analyzed, then puts each one
348 * back exactly as it was, unset included (#860).
349 *
350 * @since 2.12.0
351 * @var string[]
352 */
353 private const POSTDATA_GLOBALS = [
354 'post', 'id', 'authordata', 'currentday', 'currentmonth',
355 'page', 'pages', 'multipage', 'more', 'numpages',
356 ];
357
358 /**
359 * Resolve the content worth analyzing for a post.
360 *
361 * @param \WP_Post $post Post being analyzed.
362 * @return string HTML/text to analyze.
363 */
364 public static function resolve(\WP_Post $post): string {
365 $raw = (string) $post->post_content;
366
367 // A page built in Gutenberg and then switched to Bricks keeps its old
368 // blocks in `post_content` forever — Bricks never clears them, and
369 // never renders them either. Resolving that first meant the stale draft
370 // beat the tree the visitor actually reads, and it did not stop at the
371 // score: the same string becomes the meta description, og:description,
372 // twitter:description and the schema description. Starting from nothing
373 // sends the resolution straight to Bricks' storage, which is where this
374 // page's words are (#651).
375 //
376 // Only for the stored path. `resolve_markup()` is also called with live
377 // editor content, and the Bricks panel's resolver reads the canvas —
378 // discarding that would replace what the author is typing with the last
379 // save.
380 if (self::bricks_supersedes_post_content((int) $post->ID)) {
381 $raw = '';
382 }
383
384 return self::resolve_markup($raw, $post);
385 }
386
387 /**
388 * Whether Bricks renders this post and throws its `post_content` away.
389 *
390 * True means anything still stored in `post_content` is invisible: it is
391 * not on the page, so it must not be scored, described or published as
392 * structured data. False covers both a post Bricks does not own and a
393 * Bricks page that puts `post_content` back with a Post Content element.
394 *
395 * @since 2.3.1
396 *
397 * @param int $post_id Post being resolved.
398 * @return bool
399 */
400 public static function bricks_supersedes_post_content(int $post_id): bool {
401 $tree = self::bricks_tree($post_id);
402
403 if (empty($tree)) {
404 return false;
405 }
406
407 foreach ($tree as $element) {
408 if (is_array($element)
409 && self::BRICKS_POST_CONTENT_ELEMENT === ($element['name'] ?? null)
410 ) {
411 return false;
412 }
413 }
414
415 return !self::bricks_tree_prints_post_content($tree);
416 }
417
418 /**
419 * Whether a Bricks tree prints the body through a dynamic-data tag.
420 *
421 * The Post Content element is not the only way back onto the page: Bricks'
422 * `{post_content}` tag renders the same thing from inside an ordinary text
423 * element, and a single-post template written that way is a common shape.
424 * Missing it would mean the post's real body is discarded everywhere —
425 * scoring, the meta/og/twitter descriptions, the schema description — for a
426 * page that is displaying it.
427 *
428 * Matched over the encoded tree rather than per setting, because the tag can
429 * sit in any string field of any element and Bricks allows modifiers after
430 * the name (`{post_content:...}`).
431 *
432 * @since 2.3.1
433 *
434 * @param array $tree Bricks element tree.
435 * @return bool
436 */
437 private static function bricks_tree_prints_post_content(array $tree): bool {
438 $encoded = wp_json_encode($tree);
439
440 return is_string($encoded) && false !== stripos($encoded, '{post_content');
441 }
442
443 /**
444 * The post's content as the visitor actually receives it.
445 *
446 * `post_content` for everything except a Bricks page that discards it, and
447 * there the Bricks tree's text. Descriptions are derived from a post's body
448 * in half a dozen places; every one of them wants this rather than the raw
449 * column (#651).
450 *
451 * @since 2.3.1
452 *
453 * @param \WP_Post $post Post being described.
454 * @return string
455 */
456 public static function visible_content(\WP_Post $post): string {
457 $superseding = self::superseding_content($post);
458
459 return '' !== $superseding ? $superseding : (string) $post->post_content;
460 }
461
462 /**
463 * Replacement body text for a post whose `post_content` does not render.
464 *
465 * Empty for every ordinary post, which is what makes this safe to call from
466 * paths that already handle excerpts their own way: they keep that handling
467 * and only a Bricks page is diverted.
468 *
469 * @since 2.3.1
470 *
471 * @param \WP_Post $post Post being described.
472 * @return string Visible body text, or '' when `post_content` is fine.
473 */
474 public static function superseding_content(\WP_Post $post): string {
475 if (!self::bricks_supersedes_post_content((int) $post->ID)) {
476 return '';
477 }
478
479 $bricks = self::from_bricks((int) $post->ID);
480
481 return self::is_blank($bricks) ? '' : $bricks;
482 }
483
484 /**
485 * Body text to derive a description from, when the usual source is wrong.
486 *
487 * A hand-written excerpt is the author's own summary and is correct however
488 * the page is built, so it yields '' here and the caller's normal
489 * `get_the_excerpt()` path keeps it. Only a Bricks page with no excerpt —
490 * where core would derive one from discarded `post_content` — gets diverted.
491 *
492 * @since 2.3.1
493 *
494 * @param \WP_Post $post Post being described.
495 * @return string Text to summarize, or '' to leave the caller's path alone.
496 */
497 public static function superseding_excerpt_source(\WP_Post $post): string {
498 if ('' !== trim((string) $post->post_excerpt)) {
499 return '';
500 }
501
502 return self::superseding_content($post);
503 }
504
505 /**
506 * The Bricks element tree that renders for a post.
507 *
508 * Public because what Bricks puts on the page is not only a scoring
509 * question: the schema graph has to know whether a Bricks element already
510 * publishes the page's FAQ before adding one of its own (#649, #650).
511 *
512 * Flat, in Bricks' own storage shape — `expand_bricks_components()`
513 * appends component definitions to the same list rather than nesting them,
514 * so one `foreach` reaches every element.
515 *
516 * @since 2.3.1
517 *
518 * @param int $post_id Post being resolved.
519 * @return array<int,mixed> Elements, or [] when Bricks renders nothing here.
520 */
521 /**
522 * The builder meta keys, for callers that need to inspect the raw storage
523 * rather than the text extracted from it.
524 *
525 * The SEO Analyzer reads these to answer "is there a ThinkRank FAQ element
526 * on this post?", which is a question about the stored tree, not about the
527 * words in it (#686).
528 *
529 * @since 2.7.0
530 * @return string[]
531 */
532 public static function builder_meta_keys(): array {
533 return self::BUILDER_META_KEYS;
534 }
535
536 public static function bricks_tree(int $post_id): array {
537 if (array_key_exists($post_id, self::$bricks_trees)) {
538 return self::$bricks_trees[$post_id];
539 }
540
541 self::$bricks_trees[$post_id] = self::resolve_bricks_tree($post_id);
542
543 return self::$bricks_trees[$post_id];
544 }
545
546 /**
547 * Discard the resolved-tree memo. Test seam.
548 *
549 * @since 2.3.1
550 * @return void
551 */
552 public static function flush_bricks_cache(): void {
553 self::$bricks_trees = [];
554 }
555
556 /**
557 * Read and resolve a post's Bricks tree, ignoring the memo.
558 *
559 * @since 2.3.1
560 *
561 * @param int $post_id Post being resolved.
562 * @return array<int,mixed>
563 */
564 private static function resolve_bricks_tree(int $post_id): array {
565 if (!self::bricks_owns_post($post_id)) {
566 return [];
567 }
568
569 $source = self::bricks_content_source($post_id);
570 if (!$source) {
571 return [];
572 }
573
574 $stored = get_post_meta($source, self::bricks_meta_key(), true);
575
576 if (is_string($stored)) {
577 $stored = '' === trim($stored) ? null : json_decode($stored, true);
578 }
579
580 if (!is_array($stored) || empty($stored)) {
581 return [];
582 }
583
584 return self::expand_bricks_components(self::bricks_render_order($stored));
585 }
586
587 /**
588 * A flat Bricks element list, in the order Bricks renders it.
589 *
590 * Bricks stores one flat list and links it with `parent` and `children`
591 * ids. `Frontend::render_data()` renders the root elements in list order
592 * and each element's children in the order of its `children` array, so a
593 * child's position in the list says nothing about where it appears on the
594 * page. Walking the list as stored put a section's contents wherever they
595 * happened to be saved (#907).
596 *
597 * Anything the walk does not reach (an orphan, a cycle) keeps its stored
598 * position after the rest, so no copy is dropped.
599 *
600 * @since 2.15.0
601 *
602 * @param array $elements Flat Bricks element list.
603 * @return array The same elements, in render order.
604 */
605 private static function bricks_render_order(array $elements): array {
606 $by_id = [];
607 foreach ($elements as $index => $element) {
608 $id = is_array($element) ? ($element['id'] ?? null) : null;
609 if (is_scalar($id) && '' !== (string) $id && !isset($by_id[(string) $id])) {
610 $by_id[(string) $id] = $index;
611 }
612 }
613
614 if (empty($by_id)) {
615 return $elements;
616 }
617
618 $ordered = [];
619 $placed = [];
620
621 $place = static function ($index) use (&$place, &$ordered, &$placed, $elements, $by_id): void {
622 if (isset($placed[$index])) {
623 return;
624 }
625
626 $placed[$index] = true;
627 $ordered[] = $elements[$index];
628
629 $children = is_array($elements[$index]) ? ($elements[$index]['children'] ?? []) : [];
630 if (!is_array($children)) {
631 return;
632 }
633
634 foreach ($children as $child_id) {
635 if (is_scalar($child_id) && isset($by_id[(string) $child_id])) {
636 $place($by_id[(string) $child_id]);
637 }
638 }
639 };
640
641 foreach ($elements as $index => $element) {
642 $parent = is_array($element) ? ($element['parent'] ?? null) : null;
643 if (empty($parent) || !is_scalar($parent) || !isset($by_id[(string) $parent])) {
644 $place($index);
645 }
646 }
647
648 foreach ($elements as $index => $element) {
649 if (!isset($placed[$index])) {
650 $ordered[] = $element;
651 }
652 }
653
654 return $ordered;
655 }
656
657 /**
658 * Resolve an arbitrary chunk of editor markup for the given post.
659 *
660 * The editor sends its live content to the scorer so an author sees their
661 * unsaved edits reflected. On a builder page that live string is the raw
662 * builder markup — the block editor hands over Divi's
663 * `<!-- wp:divi/... -->` comments verbatim, because it cannot render
664 * blocks it has no client-side registration for. Analyzed as-is it reads
665 * as zero words, which is how a Divi page could show a correct saved score
666 * while the live Content Analysis panel next to it still said
667 * "No content".
668 *
669 * Running the live string through the same chain as stored content keeps
670 * both paths honest, and falling through to the post's builder storage
671 * covers builders (Oxygen) whose editor content is empty to begin with.
672 *
673 * @since 1.23.0
674 *
675 * @param string $raw Markup to analyze.
676 * @param \WP_Post $post Post the markup belongs to.
677 * @return string Content to analyze.
678 */
679 public static function resolve_markup(string $raw, \WP_Post $post): string {
680 $content = self::render_post_content($raw, $post);
681
682 // Block markup that renders to nothing usually means the builder that
683 // owns those blocks did not register them in this context — Divi 5
684 // loads its module library lazily per-request, so in CLI, REST, admin
685 // and block-editor requests do_blocks() yields an empty string while
686 // the words sit right there in the block attributes. Read them
687 // directly.
688 if (self::is_blank($content)) {
689 $from_blocks = self::from_block_attributes($raw);
690 if (!self::is_blank($from_blocks)) {
691 $content = $from_blocks;
692 }
693 }
694
695 // Only reach for builder storage when the markup yielded nothing — a
696 // classic post must never pay for this.
697 if (self::is_blank($content)) {
698 $builder = self::from_builder_meta((int) $post->ID);
699 if (!self::is_blank($builder)) {
700 $content = $builder;
701 }
702 }
703
704 // A resolution that collapsed to nothing is worse than the raw markup.
705 if (self::is_blank($content) && !self::is_blank($raw)) {
706 $content = $raw;
707 }
708
709 /**
710 * Filter the content ThinkRank analyzes for a post.
711 *
712 * Use this to teach ThinkRank about a builder it does not know, or to
713 * override extraction for one it does.
714 *
715 * @since 1.23.0
716 *
717 * @param string $content Resolved content.
718 * @param \WP_Post $post Post being analyzed.
719 * @param string $raw Markup this resolution started from.
720 */
721 return (string) apply_filters('thinkrank_analyzable_content', $content, $post, $raw);
722 }
723
724 /**
725 * Render blocks and shortcodes found in post_content.
726 *
727 * Best-effort: a third-party block that fatals must not take the whole
728 * score down with it.
729 *
730 * Runs as the post's own context, as it would on the front end. Admin and
731 * REST requests have no current post, so a shortcode reading
732 * `get_the_ID()` got nothing, and one looping a related-posts query left
733 * the global post on the last of them: its `wp_reset_postdata()` goes back
734 * to the main query's post, and there is none. On the Classic Editor this
735 * runs after the form prints its hidden `post_ID` and before the title and
736 * editor, which then showed the related post, and Update saved it over the
737 * original (#860).
738 *
739 * @param string $raw Raw post content.
740 * @param \WP_Post $post Post the content belongs to.
741 * @return string Rendered content.
742 */
743 private static function render_post_content(string $raw, \WP_Post $post): string {
744 if ('' === trim($raw)) {
745 return '';
746 }
747
748 $content = $raw;
749 $previous = self::snapshot_post_globals();
750
751 try {
752 // phpcs:ignore WordPress.WP.GlobalVariablesOverride.Prohibited -- Made current for the render, restored in finally.
753 $GLOBALS['post'] = $post;
754
755 // Fires `the_post`, which these paths never fired before: admin,
756 // REST and cron analysis had no current post at all. That is the
757 // same signal the front-end loop sends and it is what makes
758 // `get_the_ID()` work inside a shortcode, but it is a new call on a
759 // path that runs in bulk — the word-count index resolves every post
760 // it visits — so a theme that counts views on `the_post` will count
761 // them during indexing. Accepted deliberately: without it a
762 // shortcode cannot resolve its own post, which is the bug (#860).
763 if (function_exists('setup_postdata')) {
764 setup_postdata($post);
765 }
766
767 if (function_exists('has_blocks') && function_exists('do_blocks') && has_blocks($raw)) {
768 $content = do_blocks($raw);
769 }
770
771 // Block output can itself contain shortcodes, so this runs either way.
772 if (function_exists('do_shortcode') && strpos($content, '[') !== false) {
773 $content = do_shortcode($content);
774 }
775 } catch (\Throwable $e) {
776 return $raw;
777 } finally {
778 self::restore_post_globals($previous);
779 }
780
781 return self::is_blank($content) ? $raw : $content;
782 }
783
784 /**
785 * The post globals as they are now; a global that is unset has no key.
786 *
787 * @since 2.12.0
788 *
789 * @return array<string,mixed>
790 */
791 private static function snapshot_post_globals(): array {
792 $snapshot = [];
793
794 foreach (self::POSTDATA_GLOBALS as $name) {
795 if (array_key_exists($name, $GLOBALS)) {
796 $snapshot[$name] = $GLOBALS[$name];
797 }
798 }
799
800 return $snapshot;
801 }
802
803 /**
804 * Put the post globals back as snapshot_post_globals() found them.
805 *
806 * Assigned directly rather than through `setup_postdata()`: there may have
807 * been no post to set up, and re-running it would fire `the_post` again.
808 *
809 * @since 2.12.0
810 *
811 * @param array<string,mixed> $snapshot From snapshot_post_globals().
812 * @return void
813 */
814 private static function restore_post_globals(array $snapshot): void {
815 foreach (self::POSTDATA_GLOBALS as $name) {
816 if (array_key_exists($name, $snapshot)) {
817 // phpcs:ignore WordPress.NamingConventions.PrefixAllGlobals.NonPrefixedVariableFound -- Core's own globals, put back as they were.
818 $GLOBALS[$name] = $snapshot[$name];
819 } else {
820 unset($GLOBALS[$name]);
821 }
822 }
823 }
824
825 /**
826 * Extract text from the attributes of parsed blocks.
827 *
828 * @param string $raw Raw post content containing block markup.
829 * @return string Collected text, or '' when nothing was found.
830 */
831 private static function from_block_attributes(string $raw): string {
832 if (!function_exists('parse_blocks') || !function_exists('has_blocks') || !has_blocks($raw)) {
833 return '';
834 }
835
836 try {
837 $blocks = parse_blocks($raw);
838 } catch (\Throwable $e) {
839 return '';
840 }
841
842 $attrs = [];
843 $collect = static function (array $items) use (&$collect, &$attrs): void {
844 foreach ($items as $block) {
845 if (!empty($block['attrs']) && is_array($block['attrs'])) {
846 $attrs[] = $block['attrs'];
847 }
848 if (!empty($block['innerBlocks']) && is_array($block['innerBlocks'])) {
849 $collect($block['innerBlocks']);
850 }
851 }
852 };
853 $collect($blocks);
854
855 return empty($attrs) ? '' : self::text_from_tree($attrs);
856 }
857
858 /**
859 * Everything Bricks contributes to this post's analyzable content.
860 *
861 * Bricks is the only builder here that needs more than a meta key, on
862 * three counts:
863 *
864 * - It leaves its stored tree behind when a post is switched back to the
865 * block editor, so an editor-mode gate has to run first or ThinkRank
866 * scores markup the visitor never sees — the same failure
867 * `_fl_builder_draft` was ordered against in #449.
868 * - A post's content can live on ANOTHER post. Bricks' Templates feature
869 * assigns a content template by condition, and a page using one stores
870 * nothing of its own; reading only the page's meta scores it blank
871 * while the visitor reads a full page.
872 * - Its stored text carries dynamic-data tags and internal element names
873 * that never reach the rendered page.
874 *
875 * @since 2.2.1
876 *
877 * @param int $post_id Post being resolved.
878 * @return string Extracted text, or '' when Bricks has nothing for it.
879 */
880 private static function from_bricks(int $post_id): string {
881 $tree = self::bricks_tree($post_id);
882
883 if (empty($tree)) {
884 return '';
885 }
886
887 return self::strip_bricks_dynamic_tags(
888 self::text_from_tree(self::with_bricks_heading_tags(self::without_bricks_element_labels($tree)))
889 );
890 }
891
892 /**
893 * Whether Bricks — not the block editor — renders this post.
894 *
895 * Bricks writes `bricks` or `wordpress` into its editor-mode meta as the
896 * author toggles between the two, and never clears the content it stored
897 * for the other mode. Only the `wordpress` value is disqualifying: an
898 * absent value is the normal state for a post Bricks built and never
899 * toggled. This follows Bricks' own `Helpers::render_with_bricks()`, which
900 * bails on exactly that one value.
901 *
902 * It deliberately does not match it exactly: the comparison here is
903 * case-insensitive, where Bricks' is strict. Bricks 2.3.12 only ever writes
904 * the value lowercase, so the two agree on everything Bricks itself
905 * stores; they part company only on a value some other integration wrote.
906 * The two shipping today disagree about the casing — SureRank compares
907 * against `'WordPress'`, AIOSEO against `'bricks'` — and of the two ways to
908 * be wrong about `'WordPress'`, blocking costs a score on a page that has
909 * one, while allowing scores stale content the visitor never sees, which is
910 * the failure this gate exists to prevent.
911 *
912 * @since 2.2.1
913 *
914 * @param int $post_id Post being resolved.
915 * @return bool
916 */
917 private static function bricks_owns_post(int $post_id): bool {
918 $mode = get_post_meta($post_id, self::bricks_editor_mode_key(), true);
919
920 // phpcs:ignore WordPress.WP.CapitalPDangit.MisspelledInText -- Bricks' own stored meta value, lower-cased for the comparison.
921 return !(is_string($mode) && 'wordpress' === strtolower(trim($mode)));
922 }
923
924 /**
925 * The post whose Bricks tree actually renders for this post.
926 *
927 * Usually the post itself. When it stores nothing of its own, Bricks falls
928 * back to whichever content template's conditions match, and that template
929 * is a separate post carrying the words the visitor reads.
930 *
931 * Resolution is delegated to Bricks rather than reimplemented: template
932 * conditions are a whole rules engine (post IDs, types, taxonomies,
933 * archives), and a second implementation would drift from it. Bricks
934 * answers through statics, so they are saved and restored around the call —
935 * `set_active_templates()` returns early once populated, and on a
936 * front-end request Bricks has already populated it for the page being
937 * served. Clobbering that would corrupt the render in progress.
938 *
939 * Best-effort by design: any failure returns the post's own data, which is
940 * exactly today's behaviour.
941 *
942 * @since 2.2.1
943 *
944 * @param int $post_id Post being resolved.
945 * @return int Post ID holding the Bricks tree, or 0 when there is none.
946 */
947 private static function bricks_content_source(int $post_id): int {
948 $own = get_post_meta($post_id, self::bricks_meta_key(), true);
949 if ((is_array($own) && !empty($own)) || (is_string($own) && '' !== trim($own))) {
950 return $post_id;
951 }
952
953 if (!class_exists('\\Bricks\\Database')
954 || !method_exists('\\Bricks\\Database', 'set_active_templates')
955 ) {
956 return 0;
957 }
958
959 // `set_active_templates()` writes TWO statics — `$active_templates` and,
960 // when a header template resolves, `$header_position`. Both are saved,
961 // and both are restored in `finally` rather than on the happy path: a
962 // throw part-way through (a third-party hook on
963 // `bricks/database/content_type`, `bricks/builder/data_post_id` or
964 // `bricks/active_templates` is enough) must not leave Bricks' render
965 // state holding this lookup's values. Restoring only after a clean
966 // return is what the `catch` below would otherwise skip.
967 $has_header_position = property_exists('\\Bricks\\Database', 'header_position');
968 $saved_templates = \Bricks\Database::$active_templates;
969 $saved_header_position = $has_header_position ? \Bricks\Database::$header_position : null;
970
971 try {
972 \Bricks\Database::$active_templates = [];
973 \Bricks\Database::set_active_templates($post_id);
974 $template = (int) (\Bricks\Database::$active_templates['content'] ?? 0);
975 } catch (\Throwable $e) {
976 return 0;
977 } finally {
978 \Bricks\Database::$active_templates = $saved_templates;
979 if ($has_header_position) {
980 \Bricks\Database::$header_position = $saved_header_position;
981 }
982 }
983
984 // A template that is the post itself adds nothing over the empty read
985 // above, and would otherwise recurse conceptually.
986 return $template === $post_id ? 0 : $template;
987 }
988
989 /**
990 * Splice component definitions into the tree.
991 *
992 * A Bricks component keeps its markup in the `bricks_components` option,
993 * not on the page. The page stores only an instance: an element carrying
994 * `cid` and, usually, empty `settings`. Walking the page alone therefore
995 * found no words at all, and a page built entirely from components scored
996 * blank — the same failure as a page built from a content template.
997 *
998 * Confirmed on Bricks 2.3.12: `Bricks\Frontend::render_data()` renders the
999 * component's copy from an instance this walker extracted '' from.
1000 *
1001 * The definition is read straight from the option rather than through
1002 * `Bricks\Helpers::get_component_instance()`. That helper resolves an
1003 * instance's property overrides, which would be better, but it reads
1004 * `Bricks\Database::$global_data['components']` — populated once per
1005 * request, and empty in the admin and CLI contexts where bulk scoring
1006 * runs. Refreshing it would mean writing to Bricks' live render state, the
1007 * same hazard the template resolver is careful to avoid, and gating on it
1008 * would make a page score differently in wp-admin than on the front end.
1009 * Reading the stored definition is consistent everywhere.
1010 *
1011 * The trade-off: an instance that overrides a component property is scored
1012 * with the component's authored copy rather than the override. That is the
1013 * text the component renders by default, and it is much closer than the
1014 * nothing this returned before.
1015 *
1016 * @since 2.2.1
1017 *
1018 * @param array $tree Bricks content area.
1019 * @return array Tree with component elements spliced in after each instance.
1020 */
1021 private static function expand_bricks_components(array $tree): array {
1022 $expanded = [];
1023 $open = [];
1024
1025 $walk = static function (array $elements, int $depth) use (&$walk, &$expanded, &$open): void {
1026 foreach ($elements as $element) {
1027 $expanded[] = $element;
1028
1029 if (!is_array($element) || empty($element['cid']) || !is_string($element['cid'])) {
1030 continue;
1031 }
1032
1033 $cid = $element['cid'];
1034
1035 // A component nested inside its own definition would recurse
1036 // forever; the depth cap covers deep but legitimate nesting.
1037 if (isset($open[$cid]) || $depth > 4) {
1038 continue;
1039 }
1040
1041 $children = self::bricks_component_elements($cid);
1042 if (empty($children)) {
1043 continue;
1044 }
1045
1046 // Re-entrant per branch, not per page: the guard is released
1047 // after the walk so a second instance further along the page
1048 // still expands, rather than being mistaken for recursion.
1049 //
1050 // That does NOT double the word count — `text_from_tree()`
1051 // ends in `array_unique()`, which collapses a repeated
1052 // component's copy the same way it collapses a value repeated
1053 // across responsive breakpoints. Expanding both instances is
1054 // about not silently dropping the second one's structure.
1055 $open[$cid] = true;
1056 $walk($children, $depth + 1);
1057 unset($open[$cid]);
1058 }
1059 };
1060
1061 $walk($tree, 0);
1062
1063 return $expanded;
1064 }
1065
1066 /**
1067 * The stored elements of one Bricks component.
1068 *
1069 * @since 2.2.1
1070 *
1071 * @param string $cid Component id held by an instance element.
1072 * @return array Component elements, or [] when it cannot be resolved.
1073 */
1074 private static function bricks_component_elements(string $cid): array {
1075 $components = get_option(self::bricks_constant('BRICKS_DB_COMPONENTS', self::BRICKS_COMPONENTS_OPTION), []);
1076
1077 if (!is_array($components)) {
1078 return [];
1079 }
1080
1081 foreach ($components as $component) {
1082 $component = self::as_children($component);
1083 if (null === $component) {
1084 continue;
1085 }
1086
1087 if (isset($component['id']) && $component['id'] === $cid && !empty($component['elements'])) {
1088 return is_array($component['elements']) ? self::bricks_render_order($component['elements']) : [];
1089 }
1090 }
1091
1092 return [];
1093 }
1094
1095 /**
1096 * Drop each Bricks element's internal name before the tree is walked.
1097 *
1098 * A Bricks element carries an optional top-level `label` — the nickname an
1099 * author types in the Structure panel to find it again ("Hero headline",
1100 * "CTA row"). It is builder chrome and is never rendered, but `label` is in
1101 * CONTENT_KEYS because it is real content for other builders' form fields,
1102 * so it was being counted as page copy.
1103 *
1104 * Only the element's own `label` is removed. A `label` inside `settings`
1105 * is a rendered field label and stays.
1106 *
1107 * @since 2.2.1
1108 *
1109 * @param array $tree Bricks content area.
1110 * @return array Tree with element nicknames removed.
1111 */
1112 private static function without_bricks_element_labels(array $tree): array {
1113 foreach ($tree as $index => $element) {
1114 if (is_array($element) && isset($element['id'], $element['label'])) {
1115 unset($tree[$index]['label']);
1116 }
1117 }
1118
1119 return $tree;
1120 }
1121
1122 /**
1123 * Give each Bricks Heading the tag it renders with when none is stored.
1124 *
1125 * Applied on the Bricks path only. Most other builders' text nodes carry no
1126 * tag because they are not headings, so a generic "text without a tag is a
1127 * heading" rule in heading_tag_from() would turn every paragraph into one.
1128 *
1129 * A `tag` of `custom` is left alone: the element then renders its
1130 * `customTag`, which is not necessarily a heading.
1131 *
1132 * @since 2.15.0
1133 *
1134 * @param array $tree Bricks content area.
1135 * @return array Tree with each untagged Heading's default tag filled in.
1136 */
1137 private static function with_bricks_heading_tags(array $tree): array {
1138 $default = null;
1139
1140 foreach ($tree as $index => $element) {
1141 if (!is_array($element) || self::BRICKS_HEADING_ELEMENT !== ($element['name'] ?? null)) {
1142 continue;
1143 }
1144
1145 $settings = $element['settings'] ?? [];
1146 if (!is_array($settings)) {
1147 continue;
1148 }
1149
1150 $tag = $settings['tag'] ?? '';
1151 if (is_string($tag) && '' !== trim($tag)) {
1152 continue;
1153 }
1154
1155 if (null === $default) {
1156 $default = self::bricks_default_heading_tag();
1157 }
1158
1159 $settings['tag'] = $default;
1160 $tree[$index]['settings'] = $settings;
1161 }
1162
1163 return $tree;
1164 }
1165
1166 /**
1167 * The tag Bricks gives a Heading that does not set one.
1168 *
1169 * The active theme style can change it. Bricks only loads theme styles for
1170 * a front-end render, so in admin, REST and CLI requests this is the
1171 * element's own default.
1172 *
1173 * @since 2.15.0
1174 *
1175 * @return string Heading tag, h1 to h6.
1176 */
1177 private static function bricks_default_heading_tag(): string {
1178 if (class_exists('\\Bricks\\Theme_Styles')
1179 && method_exists('\\Bricks\\Theme_Styles', 'get_setting_by_key')
1180 ) {
1181 try {
1182 $styled = \Bricks\Theme_Styles::get_setting_by_key(self::BRICKS_HEADING_ELEMENT, 'tag');
1183 } catch (\Throwable $e) {
1184 $styled = null;
1185 }
1186
1187 if (is_string($styled) && preg_match('/^h[1-6]$/i', trim($styled))) {
1188 return strtolower(trim($styled));
1189 }
1190 }
1191
1192 return self::BRICKS_HEADING_DEFAULT_TAG;
1193 }
1194
1195 /**
1196 * Give each Elementor Heading widget its default `header_size` if unstored.
1197 *
1198 * @since 2.15.0
1199 *
1200 * @param array $elements Decoded `_elementor_data`.
1201 * @return array The same tree, with untagged Heading widgets tagged.
1202 */
1203 private static function with_elementor_heading_tags(array $elements): array {
1204 foreach ($elements as $index => $element) {
1205 if (!is_array($element)) {
1206 continue;
1207 }
1208
1209 if ('heading' === ($element['widgetType'] ?? null)) {
1210 $settings = $element['settings'] ?? [];
1211 if (is_array($settings)) {
1212 $size = $settings['header_size'] ?? '';
1213 if (!is_string($size) || '' === trim($size)) {
1214 $settings['header_size'] = self::ELEMENTOR_HEADING_DEFAULT_TAG;
1215 $element['settings'] = $settings;
1216 }
1217 }
1218 }
1219
1220 if (!empty($element['elements']) && is_array($element['elements'])) {
1221 $element['elements'] = self::with_elementor_heading_tags($element['elements']);
1222 }
1223
1224 $elements[$index] = $element;
1225 }
1226
1227 return $elements;
1228 }
1229
1230 /**
1231 * Remove Bricks dynamic-data tags from extracted text.
1232 *
1233 * Bricks stores `{post_title}`, `{post_meta:price}`, `{echo:my_fn}` and the
1234 * like verbatim and resolves them when it renders. Extraction reads the
1235 * stored tree, so without this the placeholders were counted as words, and
1236 * a heading whose text is `{post_title}` reported the literal token as its
1237 * heading text.
1238 *
1239 * The pattern is deliberately narrower than Bricks' own
1240 * (`/{([\wÀ-ÖØ-öø-ÿ\-\s\.\/:\(\)...]+)}/u`), which also matches braces
1241 * containing spaces. Bricks only substitutes tags that resolve to a
1242 * registered provider and leaves anything else on the page as literal text,
1243 * so the broad pattern would delete prose the visitor can actually read.
1244 * Matching only tag-shaped tokens keeps every real sentence and still
1245 * removes every placeholder — the same trade-off SureRank makes.
1246 *
1247 * @since 2.2.1
1248 *
1249 * @param string $text Extracted text.
1250 * @return string Text with placeholders removed.
1251 */
1252 private static function strip_bricks_dynamic_tags(string $text): string {
1253 $stripped = preg_replace('/\{[a-z0-9_][a-z0-9_:\-\.]*\}/i', '', $text);
1254
1255 if (null === $stripped) {
1256 return $text;
1257 }
1258
1259 // Collapse the runs of spaces a removed tag leaves mid-sentence,
1260 // without touching the newlines that separate collected nodes.
1261 $tidied = preg_replace('/[ \t]{2,}/', ' ', $stripped);
1262
1263 return null === $tidied ? $stripped : $tidied;
1264 }
1265
1266 /**
1267 * Bricks' content-area meta key, preferring Bricks' own constant.
1268 *
1269 * @since 2.2.1
1270 *
1271 * @return string
1272 */
1273 private static function bricks_meta_key(): string {
1274 return self::bricks_constant('BRICKS_DB_PAGE_CONTENT', self::BRICKS_CONTENT_META_KEY);
1275 }
1276
1277 /**
1278 * Bricks' editor-mode meta key, preferring Bricks' own constant.
1279 *
1280 * @since 2.2.1
1281 *
1282 * @return string
1283 */
1284 private static function bricks_editor_mode_key(): string {
1285 return self::bricks_constant('BRICKS_DB_EDITOR_MODE', self::BRICKS_EDITOR_MODE_META_KEY);
1286 }
1287
1288 /**
1289 * Read one of Bricks' key-name constants, falling back to the literal.
1290 *
1291 * @since 2.2.1
1292 *
1293 * @param string $name Constant name.
1294 * @param string $fallback Key to use when the constant is unavailable.
1295 * @return string
1296 */
1297 private static function bricks_constant(string $name, string $fallback): string {
1298 if (defined($name)) {
1299 $value = constant($name);
1300 if (is_string($value) && '' !== trim($value)) {
1301 return $value;
1302 }
1303 }
1304
1305 return $fallback;
1306 }
1307
1308 /**
1309 * Pull text out of whichever builder stored this post.
1310 *
1311 * @param int $post_id Post ID.
1312 * @return string Extracted text, or '' when no builder data was found.
1313 */
1314 private static function from_builder_meta(int $post_id): string {
1315 // Bricks first: it is the only builder whose content can live on
1316 // another post, and the only one gated on an editor mode.
1317 $bricks = self::from_bricks($post_id);
1318 if (!self::is_blank($bricks)) {
1319 return $bricks;
1320 }
1321
1322 foreach (self::BUILDER_META_KEYS as $key) {
1323 if (self::is_oxygen_classic_key($key)) {
1324 // Resolved as a pair, once, at the first of its keys.
1325 if ('_ct_builder_json' !== $key) {
1326 continue;
1327 }
1328
1329 $oxygen = self::from_oxygen_classic($post_id);
1330 if (!self::is_blank($oxygen)) {
1331 return $oxygen;
1332 }
1333 continue;
1334 }
1335
1336 $stored = get_post_meta($post_id, $key, true);
1337
1338 if (is_string($stored) && '' !== trim($stored)) {
1339 $decoded = json_decode($stored, true);
1340 if ('_elementor_data' === $key && is_array($decoded)) {
1341 $decoded = self::with_elementor_heading_tags($decoded);
1342 }
1343
1344 // JSON node tree (Breakdance/Oxygen 6, Elementor).
1345 if (is_array($decoded)) {
1346 $text = self::text_from_tree(self::unwrap_tree_envelope($decoded));
1347 if (!self::is_blank($text)) {
1348 return $text;
1349 }
1350 continue;
1351 }
1352
1353 continue;
1354 }
1355
1356 // Some builders store an already-decoded tree — an array for most,
1357 // an array of objects for Beaver Builder (#449).
1358 $tree = self::as_children($stored);
1359 if (null !== $tree) {
1360 $tree = self::unwrap_tree_envelope($tree);
1361 $text = self::text_from_tree($tree);
1362 if (!self::is_blank($text)) {
1363 return $text;
1364 }
1365 }
1366 }
1367
1368 return '';
1369 }
1370
1371 /**
1372 * The node tree inside a Breakdance / Oxygen 6 storage envelope.
1373 *
1374 * Neither builder stores its tree directly. The meta value is
1375 * `{"tree_json_string": "<the tree, JSON-encoded again>"}`, so one
1376 * json_decode() yields the envelope, not the tree. Walked as a tree, the
1377 * envelope is a single string leaf: kept whole as "content" when any
1378 * element held rich text (the encoded JSON then reached scoring, the
1379 * get-post-content ability and Markdown for AI), dropped when none did,
1380 * leaving the page empty (#905).
1381 *
1382 * An envelope whose inner string does not decode returns an empty tree,
1383 * never the string: handing the raw JSON back to the walker would bring
1384 * the JSON-as-content failure back on corrupt data. Anything that is not
1385 * an envelope is returned unchanged, so a bare tree still resolves.
1386 *
1387 * @since 2.15.0
1388 *
1389 * @param array $decoded Decoded meta value.
1390 * @return array The node tree.
1391 */
1392 private static function unwrap_tree_envelope(array $decoded): array {
1393 if (!array_key_exists('tree_json_string', $decoded)) {
1394 return $decoded;
1395 }
1396
1397 $inner = is_string($decoded['tree_json_string'])
1398 ? json_decode($decoded['tree_json_string'], true)
1399 : $decoded['tree_json_string'];
1400
1401 if (is_array($inner)) {
1402 return $inner;
1403 }
1404
1405 // Re-serialised envelopes can carry the tree as an object.
1406 $inner = self::as_children($inner);
1407
1408 return null !== $inner ? $inner : [];
1409 }
1410
1411 /**
1412 * Whether a meta key is one of Oxygen classic's storage keys.
1413 *
1414 * @since 2.10.0
1415 *
1416 * @param string $key Meta key.
1417 * @return bool
1418 */
1419 private static function is_oxygen_classic_key(string $key): bool {
1420 return isset(self::OXYGEN_CLASSIC_KEYS[$key]) || in_array($key, self::OXYGEN_CLASSIC_KEYS, true);
1421 }
1422
1423 /**
1424 * Text of an Oxygen classic page, from whichever stored form holds more.
1425 *
1426 * Oxygen 4.x keeps the same tree twice: as JSON, and as the shortcodes it
1427 * used before 4.0. The JSON is preferred because it carries copy the
1428 * shortcode form hides (a composite element's text is base64-encoded
1429 * inside `ct_options`, which is configuration and stripped). It is not
1430 * trusted blindly, though. Reading `ct_builder_json` first once meant a
1431 * key missing from CONTENT_KEYS silently threw the page away while the
1432 * shortcode copy sat unread next to it, because a non-empty JSON result
1433 * stopped the search. Comparing the two means the next such gap costs
1434 * nothing: the richer form wins.
1435 *
1436 * A generation is only read as a pair. The prefixed keys are what Oxygen
1437 * 4.8.3+ reads, so an unprefixed leftover next to them is stale.
1438 *
1439 * @since 2.10.0
1440 *
1441 * @param int $post_id Post ID.
1442 * @return string Extracted text, or '' when Oxygen classic stored nothing.
1443 */
1444 private static function from_oxygen_classic(int $post_id): string {
1445 foreach (self::OXYGEN_CLASSIC_KEYS as $json_key => $shortcode_key) {
1446 $json = get_post_meta($post_id, $json_key, true);
1447 $shortcodes = get_post_meta($post_id, $shortcode_key, true);
1448
1449 $from_json = '';
1450 if (is_string($json) && '' !== trim($json)) {
1451 $decoded = json_decode($json, true);
1452 if (is_array($decoded)) {
1453 // `[oxygen data="..."]` is a dynamic-data placeholder
1454 // Oxygen fills at render time. The shortcode path drops it
1455 // with every other tag, so it goes here too or the two
1456 // forms would disagree on the same page.
1457 $from_json = (string) preg_replace(
1458 '/\[oxygen\b[^\]]*\]/i',
1459 ' ',
1460 self::text_from_tree($decoded)
1461 );
1462 }
1463 }
1464
1465 $from_shortcodes = '';
1466 if (is_string($shortcodes) && strpos($shortcodes, '[') !== false) {
1467 $from_shortcodes = self::text_from_shortcodes($shortcodes);
1468 }
1469
1470 if (self::is_blank($from_json) && self::is_blank($from_shortcodes)) {
1471 continue;
1472 }
1473
1474 return self::visible_word_count($from_json) >= self::visible_word_count($from_shortcodes)
1475 ? $from_json
1476 : $from_shortcodes;
1477 }
1478
1479 return '';
1480 }
1481
1482 /**
1483 * Rough count of the words a visitor would read in extracted text.
1484 *
1485 * Only used to compare two extractions of the same page, so it needs to
1486 * be consistent rather than locale-exact.
1487 *
1488 * @since 2.10.0
1489 *
1490 * @param string $text Extracted text or markup.
1491 * @return int
1492 */
1493 private static function visible_word_count(string $text): int {
1494 $plain = trim((string) preg_replace('/\s+/u', ' ', wp_strip_all_tags($text)));
1495
1496 return '' === $plain ? 0 : count(explode(' ', $plain));
1497 }
1498
1499 /**
1500 * Shortcode attributes that carry copy a visitor reads.
1501 *
1502 * An allow-list, not a deny-list. Oxygen Classic tags carry far more
1503 * attributes than they do copy — `id`, `class`, `selector`, `url`,
1504 * `ct_options` and friends — and a deny-list silently admits every
1505 * attribute a future builder release invents, which is how markup ends up
1506 * being counted as prose.
1507 *
1508 * @var string[]
1509 */
1510 private const SHORTCODE_TEXT_ATTRIBUTES = [
1511 'text',
1512 'content',
1513 'heading',
1514 'title',
1515 'subtitle',
1516 'label',
1517 'caption',
1518 'description',
1519 'alt',
1520 'button_text',
1521 'link_text',
1522 ];
1523
1524 /**
1525 * Extract readable text from a shortcode tree, without rendering it.
1526 *
1527 * Oxygen Classic is the only builder whose storage is shortcodes rather
1528 * than JSON, and the previous implementation handed the string to
1529 * `do_shortcode()`. That silently depends on Oxygen having registered its
1530 * `ct_*` handlers in the current request — which it has on a front-end
1531 * view, and has not during bulk analysis, the post-list column, cron or
1532 * REST/MCP. With no handlers registered `do_shortcode()` returns its input
1533 * unchanged, so the raw shortcode source was scored as if it were the
1534 * page's prose: `[ct_section`, `id="section-1"` and the rest counted toward
1535 * the word count, while the actual copy sitting in `text="..."` attributes
1536 * was never counted at all (#776).
1537 *
1538 * `strip_shortcodes()` is no help either — it also only knows registered
1539 * shortcodes, so it leaves the same text untouched.
1540 *
1541 * Reading the stored tree directly is what every other builder here already
1542 * does, and it matches the class's stated design: no render engine, no
1543 * dependency on load order, safe during a bulk run.
1544 *
1545 * Parsing unconditionally, rather than rendering when Oxygen happens to be
1546 * loaded and parsing otherwise, is deliberate. It makes the extracted text
1547 * the same in every context, so the score in the editor matches the score
1548 * from a bulk run or from MCP. The old code produced whichever of the two
1549 * the request happened to allow, which is why the same post could report
1550 * two different word counts depending on how it was asked.
1551 *
1552 * The trade-off is that rendered output (resolved images, links, anything
1553 * Oxygen pulls in from a reusable part) is no longer reflected here. For
1554 * what this text feeds — word count, content scoring, meta-description
1555 * fallbacks and schema text — that markup was never the point, and counting
1556 * it only when the builder happened to be booted was the bug.
1557 *
1558 * @since 2.10.0
1559 *
1560 * @param string $stored Raw shortcode source.
1561 * @return string Extracted text.
1562 */
1563 private static function text_from_shortcodes(string $stored): string {
1564 // Oxygen stores each element's settings as a JSON blob in `ct_options`.
1565 // It is configuration, never copy, and it contains braces and brackets
1566 // that would otherwise confuse the tag scan below, so it goes first.
1567 //
1568 // The blob is matched as a balanced JSON object, not as "up to the
1569 // next quote". Oxygen wraps it in single quotes but does not escape
1570 // an apostrophe inside it (`"nicename":"Bob's Plumbing"`), so the
1571 // quote-to-quote match stopped mid-value and the rest of the blob,
1572 // `s Plumbing"}'` and all, was left in the tag and leaked into the
1573 // text. Strings inside the object are skipped whole, so neither a quote
1574 // nor a brace inside a value can end the match early.
1575 $source = (string) preg_replace(
1576 '/\sct_options\s*=\s*\'(?<obj>\{(?:[^{}"]++|"(?:[^"\\\\]|\\\\.)*+"|(?&obj))*+\})\'/s',
1577 '',
1578 $stored
1579 );
1580
1581 // Anything not shaped like Oxygen's JSON blob keeps the old,
1582 // quote-delimited strip.
1583 $source = (string) preg_replace(
1584 '/\sct_options\s*=\s*(["\']).*?\1/s',
1585 '',
1586 $source
1587 );
1588
1589 $attributes = implode('|', array_map(
1590 static fn(string $name): string => preg_quote($name, '/'),
1591 self::SHORTCODE_TEXT_ATTRIBUTES
1592 ));
1593
1594 // Replace each shortcode tag with whatever readable copy its attributes
1595 // carry. Text between tags is left exactly where it is, so the result
1596 // keeps the page's reading order rather than hoisting all the headings
1597 // to the front.
1598 // The attribute blob is matched quote-aware rather than as "anything up
1599 // to the first `]`". Oxygen copy contains brackets often enough to
1600 // matter — "Best tools [2026]", "[Updated] our policy" — and a naive
1601 // scan ends the tag inside the `text` attribute, dropping the copy
1602 // before the bracket and leaking the stray `"]` after it into the
1603 // prose. Which is this bug's own failure mode: the wrong text scored.
1604 //
1605 // A tag name must start with a letter or underscore. `[2026]` is not a
1606 // shortcode anyone can register, and scanning it as one dropped the
1607 // year out of "Best tools [2026]".
1608 $text = (string) preg_replace_callback(
1609 '/\[\/?[a-zA-Z_][a-zA-Z0-9_-]*((?:[^\]"\']|"[^"]*"|\'[^\']*\')*)\]/',
1610 static function (array $matches) use ($attributes): string {
1611 if ('' === trim($matches[1])) {
1612 return ' ';
1613 }
1614
1615 if (!preg_match_all(
1616 '/\b(' . $attributes . ')\s*=\s*(["\'])(.*?)\2/s',
1617 $matches[1],
1618 $found,
1619 PREG_SET_ORDER
1620 )) {
1621 return ' ';
1622 }
1623
1624 $parts = [];
1625 foreach ($found as $attribute) {
1626 $value = trim($attribute[3]);
1627
1628 // An attribute holding markup or a JSON fragment is
1629 // configuration that happens to share a name with a copy
1630 // field, not something a visitor reads.
1631 if ('' === $value || preg_match('/^[\[{<]/', $value)) {
1632 continue;
1633 }
1634
1635 $parts[] = $value;
1636 }
1637
1638 return empty($parts) ? ' ' : ' ' . implode(' ', $parts) . ' ';
1639 },
1640 $source
1641 );
1642
1643 // Oxygen escapes square brackets in an element's copy before writing
1644 // it between the tags, so that "Best tools [2026]" cannot be mistaken
1645 // for a shortcode (`oxygen_vsb_filter_shortcode_content_encode()`).
1646 // Decoded only now, after the tag scan, for the same reason; left
1647 // encoded, the placeholders were scored as words of their own.
1648 $text = str_replace(
1649 ['_OXY_OPENING_BRACKET_', '_OXY_CLOSING_BRACKET_'],
1650 ['[', ']'],
1651 $text
1652 );
1653
1654 // Entities are stored encoded in attributes (&amp;, &#8217;), and would
1655 // otherwise be counted as words.
1656 $text = html_entity_decode($text, ENT_QUOTES | ENT_HTML5, 'UTF-8');
1657
1658 return trim((string) preg_replace('/\s+/u', ' ', $text));
1659 }
1660
1661 /**
1662 * A node's children, whether it stores them as an array or an object.
1663 *
1664 * The walker used to return immediately on `!is_array($node)`, so an
1665 * object node was dropped along with its entire subtree — silently, as
1666 * `''`, which the caller reads as "this builder stored nothing" rather
1667 * than "this walker cannot read this shape".
1668 *
1669 * Beaver Builder stores `_fl_builder_data` as an array of stdClass nodes,
1670 * each with a stdClass `settings` object, so every node would have been
1671 * dropped and adding its meta key alone would have looked like it worked
1672 * and changed nothing. Not BB-specific: any builder storing objects hits
1673 * this, and that shape will come up again (#449).
1674 *
1675 * @since 2.1.0
1676 *
1677 * @param mixed $node Candidate node.
1678 * @return array<string|int,mixed>|null Traversable children, or null.
1679 */
1680 private static function as_children($node): ?array {
1681 if (is_array($node)) {
1682 return $node;
1683 }
1684
1685 // Deliberately not is_object(): a builder can store a value object
1686 // (DateTime, a WP_Post) whose properties are not content, and
1687 // get_object_vars() on those yields noise. stdClass is what the
1688 // JSON/serialize round-trip produces, which is the shape we want.
1689 if ($node instanceof \stdClass) {
1690 return get_object_vars($node);
1691 }
1692
1693 return null;
1694 }
1695
1696 /**
1697 * Walk a builder node tree and collect the user-visible text.
1698 *
1699 * Values are joined with block-level markup so downstream heading, link and
1700 * image detection keeps working on the result.
1701 *
1702 * One depth-first walk, so the output follows the tree's own order, which
1703 * is the order the builders read here render in. This used to be two
1704 * passes over the whole tree, one for the reconstructed headings, links and
1705 * images and one for the remaining text, and the output followed pass
1706 * order: every heading and button on the page first, every paragraph after
1707 * them. That order became the meta description, og:description, the schema
1708 * description and Pro's Markdown for AI document (#907).
1709 *
1710 * @param array $tree Decoded builder tree.
1711 * @return string Collected HTML.
1712 */
1713 private static function text_from_tree(array $tree): string {
1714 // Each entry is [value, is_markup], in tree order.
1715 $entries = [];
1716
1717 // Strings already represented inside reconstructed markup, so the plain
1718 // text doesn't emit a link label or heading a second time and double it
1719 // in the word count. Applied after the walk, against the whole tree:
1720 // a string folded into markup anywhere is dropped everywhere, exactly
1721 // as it was when the markup pass ran over the whole tree first. Checking
1722 // it during the walk instead would let a bare copy that appears before
1723 // its heading through.
1724 $consumed = [];
1725
1726 // Markup is content wherever it appears; bare strings only count when
1727 // their key says they are content, so slugs and class names stay out of
1728 // the word count.
1729 $leaf = static function ($value, $key) use (&$entries): void {
1730 if (!is_string($value) || '' === trim($value)) {
1731 return;
1732 }
1733
1734 $is_content_key = is_string($key)
1735 && in_array(strtolower($key), self::CONTENT_KEYS, true);
1736
1737 if ($is_content_key || strpos($value, '<') !== false) {
1738 $entries[] = [$value, false];
1739 }
1740 };
1741
1742 // Text only, no reconstruction. Used for a `link` / `image` / video
1743 // sub-object: a destination descriptor the parent has already folded
1744 // into its markup. Rebuilding inside it would emit the same URL a second
1745 // time as a bare link and turn an image's own `url` field into a
1746 // spurious <a>, but any copy it carries still counts.
1747 $sweep = static function ($node, $key) use (&$sweep, $leaf): void {
1748 $children = self::as_children($node);
1749 if (null === $children) {
1750 $leaf($node, $key);
1751 return;
1752 }
1753
1754 foreach ($children as $child_key => $child) {
1755 $sweep($child, is_string($child_key) ? $child_key : $key);
1756 }
1757 };
1758
1759 // Rebuild <a>, <img> and <hN> from node *shape*, then carry on through
1760 // the node's own fields in order. This has to happen per node rather
1761 // than per leaf: a link's label and its destination are separate
1762 // sibling fields, so once the tree is flattened to leaves the pairing
1763 // is gone.
1764 $walk = static function ($node, $key = null) use (&$walk, $sweep, $leaf, &$entries, &$consumed): void {
1765 $children = self::as_children($node);
1766 if (null === $children) {
1767 $leaf($node, $key);
1768 return;
1769 }
1770
1771 $markup = self::markup_for_node($children, $consumed);
1772 if ('' !== $markup) {
1773 $entries[] = [$markup, true];
1774 }
1775
1776 foreach ($children as $child_key => $child) {
1777 $next_key = is_string($child_key) ? $child_key : $key;
1778
1779 if (is_string($child_key)
1780 && (in_array(strtolower($child_key), self::URL_KEYS, true)
1781 || in_array(strtolower($child_key), self::IMAGE_KEYS, true)
1782 || in_array(strtolower($child_key), self::VIDEO_KEYS, true))
1783 ) {
1784 $sweep($child, $next_key);
1785 continue;
1786 }
1787
1788 $walk($child, $next_key);
1789 }
1790 };
1791
1792 $walk($tree);
1793
1794 $collected = [];
1795 foreach ($entries as [$value, $is_markup]) {
1796 // Already inside a reconstructed tag.
1797 if (!$is_markup && in_array($value, $consumed, true)) {
1798 continue;
1799 }
1800
1801 $collected[] = $value;
1802 }
1803
1804 if (empty($collected)) {
1805 return '';
1806 }
1807
1808 // De-duplicate: builder trees often repeat a value across responsive
1809 // breakpoints, which would otherwise multiply the word count. Keeps the
1810 // first occurrence, so a repeat never moves a value later in the page.
1811 $collected = array_unique($collected);
1812
1813 return implode("\n", $collected);
1814 }
1815
1816 /**
1817 * The video source a node is actually playing, if any.
1818 *
1819 * @since 2.3.1
1820 *
1821 * @param array $node Builder node.
1822 * @return string Video source, or '' when the node carries none.
1823 */
1824 private static function video_from(array $node): string {
1825 foreach ($node as $key => $value) {
1826 if (!is_string($key) || !is_string($value)) {
1827 continue;
1828 }
1829
1830 if (!in_array(strtolower($key), self::VIDEO_TYPE_KEYS, true)) {
1831 continue;
1832 }
1833
1834 $keys = self::VIDEO_KEYS_BY_TYPE[strtolower(trim($value))] ?? null;
1835 if (null === $keys) {
1836 continue;
1837 }
1838
1839 // A recognised video_type settles it, including when that
1840 // provider's own field is empty. Falling through to the flat sweep
1841 // there handed back whichever sibling key happened to come first in
1842 // node order — the stale youtube_url left behind after switching
1843 // the widget to a hosted file, which is exactly what keying on the
1844 // declared type is meant to prevent.
1845 $declared = self::url_from($node, $keys);
1846
1847 return self::is_video_source($declared) ? $declared : '';
1848 }
1849
1850 $url = self::url_from($node, self::VIDEO_KEYS);
1851
1852 return self::is_video_source($url) ? $url : '';
1853 }
1854
1855 /**
1856 * Whether a value can be a video source.
1857 *
1858 * `looks_like_url()` also accepts `#anchor`, `mailto:` and `tel:`, which a
1859 * link node may legitimately hold but a video cannot: `<iframe src="#top">`
1860 * is not a video and would reach a video sitemap as one.
1861 *
1862 * @since 2.3.1
1863 *
1864 * @param string $url Candidate source.
1865 * @return bool
1866 */
1867 private static function is_video_source(string $url): bool {
1868 return '' !== $url
1869 && (1 === preg_match('#^(https?:)?//#i', $url) || str_starts_with($url, '/'));
1870 }
1871
1872 /**
1873 * Whether a video source points at a file rather than a provider page.
1874 *
1875 * @since 2.3.1
1876 *
1877 * @param string $url Video source.
1878 * @return bool
1879 */
1880 private static function is_video_file(string $url): bool {
1881 $path = (string) wp_parse_url($url, PHP_URL_PATH);
1882 $ext = strtolower((string) pathinfo($path, PATHINFO_EXTENSION));
1883
1884 return in_array($ext, self::VIDEO_FILE_EXTENSIONS, true);
1885 }
1886
1887 /**
1888 * Rebuild the HTML a single builder node represents, if any.
1889 *
1890 * Looks only at the node's own fields (plus one level of nesting, because
1891 * builders commonly wrap a destination as `{ url: … }`). Returns an empty
1892 * string for the vast majority of nodes, which are layout or configuration.
1893 *
1894 * Any leaf string folded into the returned markup is appended to $consumed
1895 * so the plain-text sweep doesn't count it twice.
1896 *
1897 * @param array $node Builder node.
1898 * @param array $consumed Collects strings represented in the returned markup.
1899 * @return string Reconstructed HTML, or '' when the node carries none.
1900 */
1901 private static function markup_for_node(array $node, array &$consumed): string {
1902 $text = self::first_value($node, self::CONTENT_KEYS);
1903 $url = self::url_from($node, self::URL_KEYS);
1904 $image = self::image_from($node);
1905 $video = self::video_from($node);
1906 $tag = self::heading_tag_from($node);
1907
1908 $parts = [];
1909
1910 // Video: an embed shape rather than a link, so the video detector can
1911 // see it while the link counters do not mistake it for an outbound
1912 // link. A file source becomes <video src>, anything else an <iframe>,
1913 // matching how the builder itself renders the two cases.
1914 if ('' !== $video) {
1915 $parts[] = self::is_video_file($video)
1916 ? sprintf('<video src="%s"></video>', esc_url_raw($video))
1917 : sprintf('<iframe src="%s"></iframe>', esc_url_raw($video));
1918 }
1919
1920 // Image: alt text matters as much as the tag, since alt checks run over
1921 // whatever this returns.
1922 if ('' !== $image['url']) {
1923 $alt = '' !== $image['alt'] ? $image['alt'] : (string) self::first_value($node, self::ALT_KEYS);
1924 if ('' !== $alt) {
1925 $consumed[] = $alt;
1926 }
1927 $parts[] = sprintf(
1928 '<img src="%s" alt="%s" />',
1929 esc_url_raw($image['url']),
1930 htmlspecialchars($alt, ENT_QUOTES)
1931 );
1932 }
1933
1934 if ('' !== $text) {
1935 $inner = $text;
1936
1937 if ('' !== $url) {
1938 $consumed[] = $text;
1939 $inner = sprintf('<a href="%s">%s</a>', esc_url_raw($url), $text);
1940 }
1941
1942 if ('' !== $tag) {
1943 $consumed[] = $text;
1944 $parts[] = sprintf('<%1$s>%2$s</%1$s>', $tag, $inner);
1945 } elseif ('' !== $url) {
1946 $parts[] = $inner;
1947 }
1948 } elseif ('' !== $url) {
1949 // A destination with no label still counts as a link for link
1950 // checks; the URL doubles as its anchor text.
1951 $parts[] = sprintf('<a href="%1$s">%1$s</a>', esc_url_raw($url));
1952 }
1953
1954 return implode("\n", $parts);
1955 }
1956
1957 /**
1958 * First non-empty scalar value under any of the given keys.
1959 *
1960 * @param array $node Builder node.
1961 * @param string[] $keys Candidate keys.
1962 * @return string Trimmed value, or '' when none match.
1963 */
1964 private static function first_value(array $node, array $keys): string {
1965 foreach ($node as $key => $value) {
1966 if (!is_string($key) || !is_string($value)) {
1967 continue;
1968 }
1969 if (in_array(strtolower($key), $keys, true) && '' !== trim($value)) {
1970 return trim($value);
1971 }
1972 }
1973
1974 return '';
1975 }
1976
1977 /**
1978 * Link destination held by a node, as a bare string or a `{ url: … }` object.
1979 *
1980 * @param array $node Builder node.
1981 * @param string[] $keys Candidate keys.
1982 * @return string URL, or '' when the node holds none.
1983 */
1984 private static function url_from(array $node, array $keys): string {
1985 foreach ($node as $key => $value) {
1986 if (!is_string($key) || !in_array(strtolower($key), $keys, true)) {
1987 continue;
1988 }
1989
1990 if (is_string($value) && self::looks_like_url($value)) {
1991 return trim($value);
1992 }
1993
1994 // Elementor and Breakdance both nest the destination one level down.
1995 $nested_values = self::as_children($value);
1996 if (null !== $nested_values) {
1997 foreach ($nested_values as $nested_key => $nested) {
1998 if (is_string($nested_key)
1999 && in_array(strtolower($nested_key), ['url', 'href', 'permalink'], true)
2000 && is_string($nested)
2001 && self::looks_like_url($nested)
2002 ) {
2003 return trim($nested);
2004 }
2005 }
2006 }
2007 }
2008
2009 return '';
2010 }
2011
2012 /**
2013 * Image URL and alt text held by a node.
2014 *
2015 * @param array $node Builder node.
2016 * @return array{url:string,alt:string}
2017 */
2018 private static function image_from(array $node): array {
2019 foreach ($node as $key => $value) {
2020 if (!is_string($key) || !in_array(strtolower($key), self::IMAGE_KEYS, true)) {
2021 continue;
2022 }
2023
2024 if (is_string($value) && self::looks_like_url($value)) {
2025 return ['url' => trim($value), 'alt' => ''];
2026 }
2027
2028 $nested_values = self::as_children($value);
2029 if (null !== $nested_values) {
2030 $url = '';
2031 $alt = '';
2032 foreach ($nested_values as $nested_key => $nested) {
2033 if (!is_string($nested_key) || !is_string($nested)) {
2034 continue;
2035 }
2036 $nested_key = strtolower($nested_key);
2037 if ('' === $url && in_array($nested_key, ['url', 'src'], true) && self::looks_like_url($nested)) {
2038 $url = trim($nested);
2039 }
2040 if ('' === $alt && in_array($nested_key, self::ALT_KEYS, true)) {
2041 $alt = trim($nested);
2042 }
2043 }
2044 if ('' !== $url) {
2045 return ['url' => $url, 'alt' => $alt];
2046 }
2047 }
2048 }
2049
2050 return ['url' => '', 'alt' => ''];
2051 }
2052
2053 /**
2054 * Heading tag a node asks for, normalised to h1–h6.
2055 *
2056 * Accepts both the `h2` form and a bare level like `2`.
2057 *
2058 * @param array $node Builder node.
2059 * @return string Tag name, or '' when the node is not a heading.
2060 */
2061 private static function heading_tag_from(array $node): string {
2062 foreach ($node as $key => $value) {
2063 if (!is_string($key) || !in_array(strtolower($key), self::HEADING_TAG_KEYS, true)) {
2064 continue;
2065 }
2066
2067 if (is_string($value) && preg_match('/^h([1-6])$/i', trim($value), $m)) {
2068 return 'h' . $m[1];
2069 }
2070
2071 // A bare level only counts under a key that unambiguously means one;
2072 // `size` and `tag` carry values like "large" or "div" far more often.
2073 if (is_numeric($value)
2074 && in_array(strtolower($key), ['level'], true)
2075 && (int) $value >= 1 && (int) $value <= 6
2076 ) {
2077 return 'h' . (int) $value;
2078 }
2079 }
2080
2081 return '';
2082 }
2083
2084 /**
2085 * Whether a string is plausibly a link or asset destination.
2086 *
2087 * Deliberately permissive about relative paths — builders store internal
2088 * links that way — but rejects the option slugs and CSS values that make up
2089 * most of a builder tree.
2090 *
2091 * @param string $value Candidate.
2092 * @return bool
2093 */
2094 private static function looks_like_url(string $value): bool {
2095 $value = trim($value);
2096
2097 if ('' === $value || strlen($value) > 2048) {
2098 return false;
2099 }
2100
2101 if (preg_match('#^(https?:)?//#i', $value) || str_starts_with($value, '/')) {
2102 return true;
2103 }
2104
2105 // Protocol-ish destinations a link node can legitimately hold.
2106 return (bool) preg_match('#^(mailto:|tel:|\#)#i', $value);
2107 }
2108
2109 /**
2110 * Whether a value carries nothing worth analyzing.
2111 *
2112 * Readable text is the usual signal, but not the only one: a page can be
2113 * made entirely of media. A builder section holding just a gallery
2114 * reconstructs to `<img>` tags and one holding just a video widget to a
2115 * single `<iframe>` — both strip to an empty string, so a text-only test
2116 * discarded them here and the page fell through to the next builder key,
2117 * then to the raw markup, and finally reported as having no content at all.
2118 *
2119 * Comments are dropped before the tag test: the raw markup this class falls
2120 * back to on a builder page is unrendered block comments, which must stay
2121 * blank rather than be mistaken for reconstructed media.
2122 *
2123 * @param string $value Candidate content.
2124 * @return bool
2125 */
2126 private static function is_blank(string $value): bool {
2127 if ('' !== trim(wp_strip_all_tags($value))) {
2128 return false;
2129 }
2130
2131 $without_comments = (string) preg_replace('~<!--.*?-->~s', '', $value);
2132
2133 return 1 !== preg_match('~<(?:a|img|iframe|video|source)\b~i', $without_comments);
2134 }
2135 }
2136