PluginProbe ʕ •ᴥ•ʔ
Jetpack – WP Security, Backup, Speed, & Growth / 16.1-beta.2
Jetpack – WP Security, Backup, Speed, & Growth v16.1-beta.2
16.1-beta 16.1-beta.2 16.1-beta.3 16.1-a.5 16.1-a.3 16.0.1 16.1-a.1 16.0 16.0-beta 16.0-a.7 16.0-a.5 15.9.1 16.0-a.3 16.0-a.1 15.9 15.9-beta 15.9-a.7 15.9-a.5 15.9-a.3 15.9-a.1 15.8 15.8-beta 15.8-a.7 15.8-a.5 5.2.5 5.3.4 5.4.4 5.5.5 5.6.5 5.7.5 5.8.4 5.9.4 6.0.4 6.1 6.1.1 6.1.2 6.1.3 6.1.4 6.1.5 6.2 6.2.1 6.2.2 6.2.3 6.2.4 6.2.5 6.3 6.3.1 6.3.2 6.3.3 6.3.4 6.3.5 6.3.6 6.3.7 6.4 6.4.1 6.4.2 6.4.3 6.4.4 6.4.5 6.4.6 6.5 6.5.1 6.5.2 6.5.3 6.5.4 6.6 6.6.1 6.6.2 6.6.3 6.6.4 6.6.5 6.7 6.7.1 6.7.2 6.7.3 6.7.4 6.8 6.8.1 6.8.2 6.8.3 6.8.4 6.8.5 6.9 6.9.1 6.9.2 6.9.3 6.9.4 7.0 7.0.1 7.0.2 7.0.3 7.0.4 7.0.5 7.1 7.1.1 7.1.2 7.1.3 7.1.4 7.1.5 7.2 7.2.1 7.2.1.1 7.2.2 7.2.3 7.2.4 7.2.5 7.3 7.3.0.1 7.3.1 7.3.1.1 7.3.2 7.3.3 7.3.4 7.3.5 7.4 7.4.1 7.4.2 7.4.3 7.4.4 7.4.5 7.5 7.5.0.1 7.5.1 7.5.2 7.5.3 7.5.4 7.5.5 7.5.6 7.5.7 7.6 7.6.1 7.6.2 7.6.3 7.6.4 7.7 7.7.1 7.7.2 7.7.3 7.7.4 7.7.5 7.7.6 7.8 7.8.1 7.8.2 7.8.3 7.8.4 7.9 7.9.1 7.9.2 7.9.3 7.9.4 8.0 8.0.1 8.0.2 8.0.3 8.1 8.1.1 8.1.2 8.1.3 8.1.4 8.2 8.2.0.1 8.2.1 8.2.2 8.2.3 8.2.4 8.2.5 8.2.6 8.3 8.3.1 8.3.2 8.3.3 8.4 8.4.1 8.4.2 8.4.3 8.4.4 8.4.5 8.5 8.5.1 8.5.2 8.5.3 8.6 8.6.1 8.6.2 8.6.3 8.6.4 8.7 8.7.0.1 8.7.1 8.7.2 8.7.3 8.7.4 8.8 8.8.1 8.8.2 8.8.3 8.8.4 8.8.5 8.9 8.9.1 8.9.2 8.9.3 8.9.4 9.0 9.0.1 9.0.2 9.0.3 9.0.4 9.0.5 9.1 9.1.1 9.1.2 9.1.3 9.2 9.2.1 9.2.2 9.2.3 9.2.4 9.3 9.3.1 9.3.2 9.3.3 9.3.4 9.3.5 9.4 9.4.1 9.4.2 9.4.3 9.4.4 9.5 9.5.1 9.5.2 9.5.3 9.5.4 9.5.5 9.6 9.6.1 9.6.2 9.6.3 9.6.4 9.7 9.7.1 9.7.2 15.7-beta.2 9.7.3 15.7.1 9.8 15.8-a.1 9.8.1 15.8-a.3 9.8.2 2.0.9 9.8.3 2.1.7 9.9 2.2.10 9.9.1 2.3.10 9.9.2 2.4.7 9.9.3 2.5.5 2.6.6 2.7.5 2.8.5 2.9.6 3.0.6 3.1.5 3.2.5 3.3.6 3.4.6 3.5.6 3.6.4 3.7.5 3.8.5 3.9.10 4.0.7 4.1.4 4.2.5 4.3.5 4.4.5 4.5.3 4.6.3 4.7.4 4.8.5 4.9.3 5.0.3 5.1.4 trunk 10.0 10.0.1 10.0.2 10.1 10.1.1 10.1.2 10.2 10.2.1 10.2.2 10.2.3 10.3 10.3.1 10.3.2 10.4 10.4.1 10.4.2 10.5 10.5.1 10.5.2 10.5.3 10.6 10.6.1 10.6.2 10.7 10.7.1 10.7.2 10.8 10.8.1 10.8.2 10.9 10.9.1 10.9.2 10.9.3 11.0 11.0.1 11.0.2 11.1 11.1.1 11.1.2 11.1.3 11.1.4 11.2 11.2.1 11.2.2 11.3 11.3.1 11.3.2 11.3.3 11.3.4 11.4 11.4.1 11.4.2 11.5 11.5.1 11.5.2 11.5.3 11.6 11.6.1 11.6.2 11.7 11.7.1 11.7.2 11.7.3 11.8 11.8.3 11.8.4 11.8.5 11.8.6 11.9 11.9.1 11.9.2 11.9.3 12.0 12.0.1 12.0.2 12.1 12.1.1 12.1.2 12.2 12.2.1 12.2.2 12.3 12.3.1 12.4 12.4.1 12.5 12.5.1 12.6 12.6.1 12.6.2 12.6.3 12.7 12.7.1 12.7.2 12.8 12.8.1 12.8.2 12.9 12.9.1 12.9.2 12.9.3 12.9.4 13.0 13.0.1 13.1 13.1.1 13.1.2 13.1.3 13.1.4 13.2 13.2.1 13.2.2 13.2.3 13.3 13.3.1 13.3.2 13.4 13.4.1 13.4.2 13.4.3 13.4.4 13.5 13.5.1 13.6 13.6.1 13.7 13.7.1 13.8 13.8.1 13.8.2 13.9 13.9.1 14.0 14.1 14.2 14.2.1 14.3 14.4 14.4.1 14.5 14.6 14.7 14.8 14.9 14.9.1 15.0 15.0.1 15.0.2 15.1 15.1.1 15.2 15.3 15.3.1 15.4 15.5 15.6 15.7 15.7-a.1 15.7-a.3 15.7-a.5 15.7-a.7 15.7-beta
jetpack / vendor / wp-php-toolkit / xml / class-xmldecoder.php
jetpack / vendor / wp-php-toolkit / xml Last commit date
PHP 6 days ago LICENSE.md 6 days ago README.md 6 days ago class-xmlattributetoken.php 6 days ago class-xmldecoder.php 6 days ago class-xmlelement.php 6 days ago class-xmlnativecursorprocessor.php 6 days ago class-xmlprocessor.php 6 days ago class-xmlunsupportedexception.php 6 days ago composer.json 6 days ago
class-xmldecoder.php
227 lines
1 <?php
2
3 namespace WordPress\XML;
4
5 use function WordPress\Encoding\codepoint_to_utf8_bytes;
6
7 /**
8 * XML API: WP_XML_Decoder class
9 *
10 * Decodes spans of raw text found inside XML content,
11 * whether found in an attribute or in a text node.
12 *
13 * Do not use this function on the contents of a CDATA section,
14 * as those sections are not encoded with the XML rules unless
15 * they are embedded XML content.
16 *
17 * @package WordPress
18 * @subpackage HTML-API
19 * @since WP_VERSION
20 */
21 class XMLDecoder {
22 /**
23 * Decodes a span of XML text.
24 *
25 * Example:
26 *
27 * '&' = WP_XML_Decoder::decode( '&amp;' );
28 * '…' = WP_XML_Decoder::decode( '&#x2026;' );
29 *
30 * @todo Add examples of parse failures, and decide if it should fail or not.
31 *
32 * @since WP_VERSION
33 *
34 * @access private
35 *
36 * @param string $text Text document containing span of text to decode.
37 * @return string Decoded UTF-8 string.
38 */
39 public static function decode( $text ) {
40 $decoded = '';
41 $end = strlen( $text );
42 $at = 0;
43 $was_at = 0;
44
45 while ( $at < $end ) {
46 $next_character_reference_at = strpos( $text, '&', $at );
47 if ( false === $next_character_reference_at || $next_character_reference_at >= $end ) {
48 break;
49 }
50
51 $start_of_potential_reference_at = $next_character_reference_at + 1;
52 if ( $start_of_potential_reference_at >= $end ) {
53 // @todo This is an error. The document ended too early; consume the rest as plaintext, which is wrong.
54 break;
55 }
56
57 /**
58 * First character after the opening `&`.
59 */
60 $start_of_potential_reference = $text[ $start_of_potential_reference_at ];
61
62 /*
63 * If it's a named character reference, it will be one of the five mandated references.
64 *
65 * - `&amp;`
66 * - `&apos;`
67 * - `&gt;`
68 * - `&lt;`
69 * - `&quot;`
70 *
71 * These all must be found within the five successive characters from the `&`.
72 *
73 * Example:
74 *
75 * ╭ ampersand at 9 = $end - 6
76 * &apos;XML&apos; ($end = 15)
77 * ╰───┴─ this length must be at least 5 long,
78 * which is $end - 5.
79 */
80 if (
81 $next_character_reference_at < $end - 5 &&
82 (
83 'a' === $start_of_potential_reference ||
84 'g' === $start_of_potential_reference ||
85 'l' === $start_of_potential_reference ||
86 'q' === $start_of_potential_reference
87 )
88 ) {
89 foreach ( array(
90 'amp;' => '&',
91 'apos;' => "'",
92 'lt;' => '<',
93 'gt;' => '>',
94 'quot;' => '"',
95 ) as $name => $substitution ) {
96 if ( 0 === substr_compare( $text, $name, $start_of_potential_reference_at, strlen( $name ) ) ) {
97 $decoded .= substr( $text, $was_at, $next_character_reference_at - $was_at ) . $substitution;
98 $at = $start_of_potential_reference_at + strlen( $name );
99 $was_at = $at;
100 continue 2;
101 }
102 }
103
104 // @todo This is an invalid document. It should be communicated. Treat as plaintext and continue.
105 ++$at;
106 continue;
107 }
108
109 /*
110 * The shortest numerical character reference is four characters.
111 *
112 * Example:
113 *
114 * &#9;
115 */
116 if ( '#' !== $start_of_potential_reference || $next_character_reference_at + 4 >= $end ) {
117 // @todo This is an error. This ampersand _must_ be encoded. Treat as plaintext and move on.
118 ++$at;
119 continue;
120 }
121
122 $is_hex = 'x' === $text[ $start_of_potential_reference_at + 1 ];
123 if ( $is_hex ) {
124 $zeros_at = $start_of_potential_reference_at + 2;
125 $base = 16;
126 $digit_chars = '0123456789abcdefABCDEF';
127 $max_digits = 6; // `&#x10FFFF;`.
128 } else {
129 $zeros_at = $start_of_potential_reference_at + 1;
130 $base = 10;
131 $digit_chars = '0123456789';
132 $max_digits = 7; // `&#1114111;`.
133 }
134
135 $zero_count = strspn( $text, '0', $zeros_at );
136 $digits_at = $zeros_at + $zero_count;
137 $digit_count = strspn( $text, $digit_chars, $digits_at, $max_digits );
138 $semi_at = $digits_at + $digit_count;
139
140 if ( 0 === $digit_count || $semi_at >= $end || ';' !== $text[ $semi_at ] ) {
141 // @todo This is an error. Treat as plaintext and move on.
142 ++$at;
143 continue;
144 }
145
146 $codepoint = intval( substr( $text, $digits_at, $digit_count ), $base );
147 $character_reference = codepoint_to_utf8_bytes( $codepoint );
148 if ( '' === $character_reference && 0xFFFD !== $codepoint ) {
149 /*
150 * Stop processing if we got an invalid character AND the reference does not
151 * specifically refer code point FFFD (�).
152 *
153 * > It is a fatal error when an XML processor encounters an entity with an
154 * > encoding that it is unable to process. It is a fatal error if an XML entity
155 * > is determined (via default, encoding declaration, or higher-level protocol)
156 * > to be in a certain encoding but contains byte sequences that are not legal
157 * > in that encoding. Specifically, it is a fatal error if an entity encoded in
158 * > UTF-8 contains any ill-formed code unit sequences, as defined in section
159 * > 3.9 of Unicode [Unicode]. Unless an encoding is determined by a higher-level
160 * > protocol, it is also a fatal error if an XML entity contains no encoding
161 * > declaration and its content is not legal UTF-8 or UTF-16.
162 *
163 * See https://www.w3.org/TR/xml/#charencoding
164 */
165 // @todo This is an error. Treat as plaintext and continue, which is wrong.
166 ++$at;
167 continue;
168 }
169
170 $decoded .= substr( $text, $was_at, $at - $was_at );
171 $decoded .= $character_reference;
172 $at = $semi_at + 1;
173 $was_at = $at;
174 }
175
176 if ( 0 === $was_at ) {
177 return $text;
178 }
179
180 if ( $was_at < $end ) {
181 $decoded .= substr( $text, $was_at, $end - $was_at );
182 }
183
184 return $decoded;
185 }
186
187 /**
188 * Finds and parses the next entity in a given text starting after the
189 * given byte offset, and being entirely found within the given max length.
190 *
191 * @since {WP_VERSION}
192 *
193 * // @todo Implement this function.
194 *
195 * @param string $text Text in which to search for an XML entity.
196 * @param int $starting_byte_offset Start looking after this byte offset.
197 * @param int $ending_byte_offset Stop looking if entity is not fully contained before this byte offset.
198 * @param int|null $entity_at Optional. If provided, will be set to byte offset where entity was
199 * found, if found. Otherwise, will not be set.
200 *
201 * @return string|null Parsed entity, if parsed, otherwise `null`.
202 */
203 public static function next_entity( string $text, int $starting_byte_offset, int $ending_byte_offset, ?int &$entity_at = null ): ?string {
204 $at = $starting_byte_offset;
205 $end = $ending_byte_offset;
206
207 while ( $at < $end ) {
208 $remaining = $end - $at;
209 $amp_after = strcspn( $text, '&', $at, $remaining );
210
211 // There are no more possible entities.
212 if ( $amp_after === $remaining ) {
213 return null;
214 }
215
216 /*
217 * @todo Move the decoding logic from `decode()` above into here,
218 * then call this function in a loop from `decode()`.
219 */
220
221 ++$at;
222 }
223
224 return null;
225 }
226 }
227