PluginProbe ʕ •ᴥ•ʔ
Jetpack – WP Security, Backup, Speed, & Growth / 16.1-beta.3
Jetpack – WP Security, Backup, Speed, & Growth v16.1-beta.3
16.1-beta 16.1-beta.2 16.1-beta.3 16.1-a.5 16.1-a.3 16.0.1 16.1-a.1 16.0 16.0-beta 16.0-a.7 16.0-a.5 15.9.1 16.0-a.3 16.0-a.1 15.9 15.9-beta 15.9-a.7 15.9-a.5 15.9-a.3 15.9-a.1 15.8 15.8-beta 15.8-a.7 15.8-a.5 5.2.5 5.3.4 5.4.4 5.5.5 5.6.5 5.7.5 5.8.4 5.9.4 6.0.4 6.1 6.1.1 6.1.2 6.1.3 6.1.4 6.1.5 6.2 6.2.1 6.2.2 6.2.3 6.2.4 6.2.5 6.3 6.3.1 6.3.2 6.3.3 6.3.4 6.3.5 6.3.6 6.3.7 6.4 6.4.1 6.4.2 6.4.3 6.4.4 6.4.5 6.4.6 6.5 6.5.1 6.5.2 6.5.3 6.5.4 6.6 6.6.1 6.6.2 6.6.3 6.6.4 6.6.5 6.7 6.7.1 6.7.2 6.7.3 6.7.4 6.8 6.8.1 6.8.2 6.8.3 6.8.4 6.8.5 6.9 6.9.1 6.9.2 6.9.3 6.9.4 7.0 7.0.1 7.0.2 7.0.3 7.0.4 7.0.5 7.1 7.1.1 7.1.2 7.1.3 7.1.4 7.1.5 7.2 7.2.1 7.2.1.1 7.2.2 7.2.3 7.2.4 7.2.5 7.3 7.3.0.1 7.3.1 7.3.1.1 7.3.2 7.3.3 7.3.4 7.3.5 7.4 7.4.1 7.4.2 7.4.3 7.4.4 7.4.5 7.5 7.5.0.1 7.5.1 7.5.2 7.5.3 7.5.4 7.5.5 7.5.6 7.5.7 7.6 7.6.1 7.6.2 7.6.3 7.6.4 7.7 7.7.1 7.7.2 7.7.3 7.7.4 7.7.5 7.7.6 7.8 7.8.1 7.8.2 7.8.3 7.8.4 7.9 7.9.1 7.9.2 7.9.3 7.9.4 8.0 8.0.1 8.0.2 8.0.3 8.1 8.1.1 8.1.2 8.1.3 8.1.4 8.2 8.2.0.1 8.2.1 8.2.2 8.2.3 8.2.4 8.2.5 8.2.6 8.3 8.3.1 8.3.2 8.3.3 8.4 8.4.1 8.4.2 8.4.3 8.4.4 8.4.5 8.5 8.5.1 8.5.2 8.5.3 8.6 8.6.1 8.6.2 8.6.3 8.6.4 8.7 8.7.0.1 8.7.1 8.7.2 8.7.3 8.7.4 8.8 8.8.1 8.8.2 8.8.3 8.8.4 8.8.5 8.9 8.9.1 8.9.2 8.9.3 8.9.4 9.0 9.0.1 9.0.2 9.0.3 9.0.4 9.0.5 9.1 9.1.1 9.1.2 9.1.3 9.2 9.2.1 9.2.2 9.2.3 9.2.4 9.3 9.3.1 9.3.2 9.3.3 9.3.4 9.3.5 9.4 9.4.1 9.4.2 9.4.3 9.4.4 9.5 9.5.1 9.5.2 9.5.3 9.5.4 9.5.5 9.6 9.6.1 9.6.2 9.6.3 9.6.4 9.7 9.7.1 9.7.2 15.7-beta.2 9.7.3 15.7.1 9.8 15.8-a.1 9.8.1 15.8-a.3 9.8.2 2.0.9 9.8.3 2.1.7 9.9 2.2.10 9.9.1 2.3.10 9.9.2 2.4.7 9.9.3 2.5.5 2.6.6 2.7.5 2.8.5 2.9.6 3.0.6 3.1.5 3.2.5 3.3.6 3.4.6 3.5.6 3.6.4 3.7.5 3.8.5 3.9.10 4.0.7 4.1.4 4.2.5 4.3.5 4.4.5 4.5.3 4.6.3 4.7.4 4.8.5 4.9.3 5.0.3 5.1.4 trunk 10.0 10.0.1 10.0.2 10.1 10.1.1 10.1.2 10.2 10.2.1 10.2.2 10.2.3 10.3 10.3.1 10.3.2 10.4 10.4.1 10.4.2 10.5 10.5.1 10.5.2 10.5.3 10.6 10.6.1 10.6.2 10.7 10.7.1 10.7.2 10.8 10.8.1 10.8.2 10.9 10.9.1 10.9.2 10.9.3 11.0 11.0.1 11.0.2 11.1 11.1.1 11.1.2 11.1.3 11.1.4 11.2 11.2.1 11.2.2 11.3 11.3.1 11.3.2 11.3.3 11.3.4 11.4 11.4.1 11.4.2 11.5 11.5.1 11.5.2 11.5.3 11.6 11.6.1 11.6.2 11.7 11.7.1 11.7.2 11.7.3 11.8 11.8.3 11.8.4 11.8.5 11.8.6 11.9 11.9.1 11.9.2 11.9.3 12.0 12.0.1 12.0.2 12.1 12.1.1 12.1.2 12.2 12.2.1 12.2.2 12.3 12.3.1 12.4 12.4.1 12.5 12.5.1 12.6 12.6.1 12.6.2 12.6.3 12.7 12.7.1 12.7.2 12.8 12.8.1 12.8.2 12.9 12.9.1 12.9.2 12.9.3 12.9.4 13.0 13.0.1 13.1 13.1.1 13.1.2 13.1.3 13.1.4 13.2 13.2.1 13.2.2 13.2.3 13.3 13.3.1 13.3.2 13.4 13.4.1 13.4.2 13.4.3 13.4.4 13.5 13.5.1 13.6 13.6.1 13.7 13.7.1 13.8 13.8.1 13.8.2 13.9 13.9.1 14.0 14.1 14.2 14.2.1 14.3 14.4 14.4.1 14.5 14.6 14.7 14.8 14.9 14.9.1 15.0 15.0.1 15.0.2 15.1 15.1.1 15.2 15.3 15.3.1 15.4 15.5 15.6 15.7 15.7-a.1 15.7-a.3 15.7-a.5 15.7-a.7 15.7-beta
jetpack / vendor / wp-php-toolkit / encoding / utf8.php
jetpack / vendor / wp-php-toolkit / encoding Last commit date
LICENSE.md 6 days ago README.md 6 days ago compat-utf8.php 6 days ago composer.json 6 days ago utf8-encoder.php 6 days ago utf8.php 6 days ago
utf8.php
228 lines
1 <?php
2
3 namespace WordPress\Encoding;
4
5 use function WordPress\Encoding\compat\_wp_is_valid_utf8_fallback;
6 use function WordPress\Encoding\compat\_wp_scrub_utf8_fallback;
7 use function WordPress\Encoding\compat\_wp_has_noncharacters_fallback;
8
9 if ( extension_loaded( 'mbstring' ) ) :
10 /**
11 * Determines if a given byte string represents a valid UTF-8 encoding.
12 *
13 * Note that it’s unlikely for non-UTF-8 data to validate as UTF-8, but
14 * it is still possible. Many texts are simultaneously valid UTF-8,
15 * valid US-ASCII, and valid ISO-8859-1 (`latin1`).
16 *
17 * Example:
18 *
19 * true === wp_is_valid_utf8( '' );
20 * true === wp_is_valid_utf8( 'just a test' );
21 * true === wp_is_valid_utf8( "\xE2\x9C\x8F" ); // Pencil, U+270F.
22 * true === wp_is_valid_utf8( "\u{270F}" ); // Pencil, U+270F.
23 * true === wp_is_valid_utf8( '✏' ); // Pencil, U+270F.
24 *
25 * false === wp_is_valid_utf8( "just \xC0 test" ); // Invalid bytes.
26 * false === wp_is_valid_utf8( "\xE2\x9C" ); // Invalid/incomplete sequences.
27 * false === wp_is_valid_utf8( "\xC1\xBF" ); // Overlong sequences.
28 * false === wp_is_valid_utf8( "\xED\xB0\x80" ); // Surrogate halves.
29 * false === wp_is_valid_utf8( "B\xFCch" ); // ISO-8859-1 high-bytes.
30 * // E.g. The “ü” in ISO-8859-1 is a single byte 0xFC,
31 * // but in UTF-8 is the two-byte sequence 0xC3 0xBC.
32 *
33 * A “valid” string consists of “well-formed UTF-8 code unit sequence[s],” meaning
34 * that the bytes conform to the UTF-8 encoding scheme, all characters use the minimal
35 * byte sequence required by UTF-8, and that no sequence encodes a UTF-16 surrogate
36 * code point or any character above the representable range.
37 *
38 * @see https://www.unicode.org/versions/Unicode16.0.0/core-spec/chapter-3/#G32860
39 *
40 * @since 6.9.0
41 *
42 * @param string $bytes String which might contain text encoded as UTF-8.
43 * @return bool Whether the provided bytes can decode as valid UTF-8.
44 */
45 function wp_is_valid_utf8( string $bytes ): bool {
46 return mb_check_encoding( $bytes, 'UTF-8' );
47 }
48 else :
49 /**
50 * Fallback function for validating UTF-8.
51 *
52 * @ignore
53 * @private
54 *
55 * @since 6.9.0
56 */
57 // phpcs:ignore Universal.NamingConventions.NoReservedKeywordParameterNames.stringFound
58 function wp_is_valid_utf8( string $string ): bool {
59 return _wp_is_valid_utf8_fallback( $string );
60 }
61 endif;
62
63 if (
64 extension_loaded( 'mbstring' ) &&
65 // Maximal subpart substitution introduced by php/php-src@04e59c916f12b322ac55f22314e31bd0176d01cb.
66 version_compare( PHP_VERSION, '8.1.6', '>=' )
67 ) :
68 /**
69 * Replaces ill-formed UTF-8 byte sequences with the Unicode Replacement Character.
70 *
71 * Knowing what to do in the presence of text encoding issues can be complicated.
72 * This function replaces invalid spans of bytes to neutralize any corruption that
73 * may be there and prevent it from causing further problems downstream.
74 *
75 * However, it’s not always ideal to replace those bytes. In some settings it may
76 * be best to leave the invalid bytes in the string so that downstream code can handle
77 * them in a specific way. Replacing the bytes too early, like escaping for HTML too
78 * early, can introduce other forms of corruption and data loss.
79 *
80 * When in doubt, use this function to replace spans of invalid bytes.
81 *
82 * Replacement follows the “maximal subpart” algorithm for secure and interoperable
83 * strings. This can lead to sequences of multiple replacement characters in a row.
84 *
85 * Example:
86 *
87 * // Valid strings come through unchanged.
88 * 'test' === wp_scrub_utf8( 'test' );
89 *
90 * // Invalid sequences of bytes are replaced.
91 * $invalid = "the byte \xC0 is never allowed in a UTF-8 string.";
92 * "the byte \u{FFFD} is never allowed in a UTF-8 string." === wp_scrub_utf8( $invalid, true );
93 * 'the byte � is never allowed in a UTF-8 string.' === wp_scrub_utf8( $invalid, true );
94 *
95 * // Maximal subparts are replaced individually.
96 * '.�.' === wp_scrub_utf8( ".\xC0." ); // C0 is never valid.
97 * '.�.' === wp_scrub_utf8( ".\xE2\x8C." ); // Missing A3 at end.
98 * '.��.' === wp_scrub_utf8( ".\xE2\x8C\xE2\x8C." ); // Maximal subparts replaced separately.
99 * '.��.' === wp_scrub_utf8( ".\xC1\xBF." ); // Overlong sequence.
100 * '.���.' === wp_scrub_utf8( ".\xED\xA0\x80." ); // Surrogate half.
101 *
102 * Note! The Unicode Replacement Character is itself a Unicode character (U+FFFD).
103 * Once a span of invalid bytes has been replaced by one, it will not be possible
104 * to know whether the replacement character was originally intended to be there
105 * or if it is the result of scrubbing bytes. It is ideal to leave replacement for
106 * display only, but some contexts (e.g. generating XML or passing data into a
107 * large language model) require valid input strings.
108 *
109 * @since 6.9.0
110 *
111 * @see https://www.unicode.org/versions/Unicode16.0.0/core-spec/chapter-5/#G40630
112 *
113 * @param string $text String which is assumed to be UTF-8 but may contain invalid sequences of bytes.
114 * @return string Input text with invalid sequences of bytes replaced with the Unicode replacement character.
115 */
116 function wp_scrub_utf8( $text ) {
117 /*
118 * While it looks like setting the substitute character could fail,
119 * the internal PHP code will never fail when provided a valid
120 * code point as a number. In this case, there’s no need to check
121 * its return value to see if it succeeded.
122 */
123 $prev_replacement_character = mb_substitute_character();
124 mb_substitute_character( 0xFFFD );
125 $scrubbed = mb_scrub( $text, 'UTF-8' );
126 mb_substitute_character( $prev_replacement_character );
127
128 return $scrubbed;
129 }
130 else :
131 /**
132 * Fallback function for scrubbing UTF-8.
133 *
134 * @ignore
135 * @private
136 *
137 * @since 6.9.0
138 */
139 function wp_scrub_utf8( $text ) {
140 return _wp_scrub_utf8_fallback( $text );
141 }
142 endif;
143
144 function _wp_can_use_pcre_u( $set = null ) {
145 static $utf8_pcre = 'reset';
146
147 if ( null !== $set ) {
148 $utf8_pcre = $set;
149 }
150
151 if ( 'reset' === $utf8_pcre ) {
152 // phpcs:ignore WordPress.PHP.NoSilencedErrors.Discouraged -- intentional error generated to detect PCRE/u support.
153 $utf8_pcre = @preg_match( '/^./u', 'a' );
154 }
155
156 return $utf8_pcre;
157 }
158
159 if ( _wp_can_use_pcre_u() ) :
160 /**
161 * Returns whether the given string contains Unicode noncharacters.
162 *
163 * XML recommends against using noncharacters and HTML forbids their
164 * use in attribute names. Unicode recommends that they not be used
165 * in open exchange of data.
166 *
167 * Noncharacters are code points within the following ranges:
168 * - U+FDD0–U+FDEF
169 * - U+FFFE–U+FFFF
170 * - U+1FFFE, U+1FFFF, U+2FFFE, U+2FFFF, …, U+10FFFE, U+10FFFF
171 *
172 * @see https://www.unicode.org/versions/Unicode17.0.0/core-spec/chapter-23/#G12612
173 * @see https://www.w3.org/TR/xml/#charsets
174 * @see https://html.spec.whatwg.org/#attributes-2
175 *
176 * @since 6.9.0
177 *
178 * @param string $text Are there noncharacters in this string?
179 * @return bool Whether noncharacters were found in the string.
180 */
181 function wp_has_noncharacters( string $text ): bool {
182 return 1 === preg_match(
183 '/[\x{FDD0}-\x{FDEF}\x{FFFE}\x{FFFF}\x{1FFFE}\x{1FFFF}\x{2FFFE}\x{2FFFF}\x{3FFFE}\x{3FFFF}\x{4FFFE}\x{4FFFF}\x{5FFFE}\x{5FFFF}\x{6FFFE}\x{6FFFF}\x{7FFFE}\x{7FFFF}\x{8FFFE}\x{8FFFF}\x{9FFFE}\x{9FFFF}\x{AFFFE}\x{AFFFF}\x{BFFFE}\x{BFFFF}\x{CFFFE}\x{CFFFF}\x{DFFFE}\x{DFFFF}\x{EFFFE}\x{EFFFF}\x{FFFFE}\x{FFFFF}\x{10FFFE}\x{10FFFF}]/u',
184 $text
185 );
186 }
187 else :
188 /**
189 * Fallback function for detecting noncharacters in a text.
190 *
191 * @ignore
192 * @private
193 *
194 * @since 6.9.0
195 */
196 function wp_has_noncharacters( string $text ): bool {
197 return _wp_has_noncharacters_fallback( $text );
198 }
199 endif;
200
201 /**
202 * Convert a UTF-8 byte sequence to its Unicode codepoint.
203 *
204 * @param string $character UTF-8 encoded byte sequence representing a single Unicode character.
205 *
206 * @return int Unicode codepoint.
207 */
208 function utf8_ord( string $character ): int {
209 // Convert the byte sequence to its binary representation.
210 $bytes = unpack( 'C*', $character );
211
212 // Initialize the codepoint.
213 $codepoint = 0;
214
215 // Calculate the codepoint based on the number of bytes.
216 if ( 1 === count( $bytes ) ) {
217 $codepoint = $bytes[1];
218 } elseif ( 2 === count( $bytes ) ) {
219 $codepoint = ( ( $bytes[1] & 0x1F ) << 6 ) | ( $bytes[2] & 0x3F );
220 } elseif ( 3 === count( $bytes ) ) {
221 $codepoint = ( ( $bytes[1] & 0x0F ) << 12 ) | ( ( $bytes[2] & 0x3F ) << 6 ) | ( $bytes[3] & 0x3F );
222 } elseif ( 4 === count( $bytes ) ) {
223 $codepoint = ( ( $bytes[1] & 0x07 ) << 18 ) | ( ( $bytes[2] & 0x3F ) << 12 ) | ( ( $bytes[3] & 0x3F ) << 6 ) | ( $bytes[4] & 0x3F );
224 }
225
226 return $codepoint;
227 }
228