| 1 |
--- |
| 2 |
title: Quick Start |
| 3 |
--- |
| 4 |
|
| 5 |
Find below sample code that demonstrate the fundamental features of PHP Simple HTML DOM Parser. |
| 6 |
|
| 7 |
## Read plain text from HTML document |
| 8 |
|
| 9 |
```php |
| 10 |
<?php |
| 11 |
include_once 'HtmlWeb.php'; |
| 12 |
use simplehtmldom\HtmlWeb; |
| 13 |
|
| 14 |
$html = new HtmlWeb(); |
| 15 |
echo $html->load('https://www.google.com/')->plaintext; |
| 16 |
``` |
| 17 |
|
| 18 |
Loads a webpage into memory, parses it and returns the plain text. |
| 19 |
|
| 20 |
## Read plain text from HTML string |
| 21 |
|
| 22 |
```php |
| 23 |
<?php |
| 24 |
include_once 'HtmlDocument.php'; |
| 25 |
use simplehtmldom\HtmlDocument; |
| 26 |
|
| 27 |
$html = new HtmlDocument(); |
| 28 |
echo $html->load('<ul><li>Hello, World!</li></ul>')->plaintext; |
| 29 |
``` |
| 30 |
|
| 31 |
Parses HTML formatted text and returns the plain text. Note that the parser handles partial documents as well as full documents. |
| 32 |
|
| 33 |
## Read specific elements from HTML document |
| 34 |
|
| 35 |
```php |
| 36 |
<?php |
| 37 |
include_once 'HtmlWeb.php'; |
| 38 |
use simplehtmldom\HtmlWeb; |
| 39 |
|
| 40 |
$html = new HtmlWeb(); |
| 41 |
$html->load('https://www.google.com/'); |
| 42 |
|
| 43 |
foreach($html->find('img') as $element) |
| 44 |
echo $element->src . '<br>'; |
| 45 |
|
| 46 |
foreach($html->find('a') as $element) |
| 47 |
echo $element->href . '<br>'; |
| 48 |
``` |
| 49 |
|
| 50 |
Loads the specified document into memory and returns a list of image sources as well as anchor links. Note that [](examples/finding-html-elements.md`find`](examples/finding-html-elements.md](examples/finding-html-elements.md) supports [](https://www.w3.org/TR/selectors/CSS](https://www.w3.org/TR/selectors/](https://www.w3.org/TR/selectors/) selectors to find elements in the DOM. |
| 51 |
|
| 52 |
## Modify HTML documents |
| 53 |
|
| 54 |
```php |
| 55 |
<?php |
| 56 |
include_once 'HtmlDocument.php'; |
| 57 |
use simplehtmldom\HtmlDocument; |
| 58 |
|
| 59 |
$html = new HtmlDocument(); |
| 60 |
$html->load('<div id="hello">Hello, </div><div id="world">World!</div>'); |
| 61 |
|
| 62 |
$html->find('div', 1)->class = 'bar'; |
| 63 |
$html->find('div[id=hello]', 0)->innertext = 'foo'; |
| 64 |
|
| 65 |
echo $html; // <div id="hello">foo</div><div id="world" class="bar">World!</div> |
| 66 |
``` |
| 67 |
|
| 68 |
Parses the provided HTML string and replaces elements in the DOM before returning the updated HTML string. In this example, the class for the second `div` element is set to `bar` and the inner text for the first `div` element to `foo`. |
| 69 |
|
| 70 |
Note that [`find`](examples/finding-html-elements.md) supports a second parameter to return a single element from the array of matches. |
| 71 |
|
| 72 |
Note that attributes can be accessed directly by the means of magic methods (`->class` and `->innertext` in the example above). |
| 73 |
|
| 74 |
## Collect information from Slashdot |
| 75 |
|
| 76 |
```php |
| 77 |
<?php |
| 78 |
include_once 'HtmlWeb.php'; |
| 79 |
use simplehtmldom\HtmlWeb; |
| 80 |
|
| 81 |
$html = new HtmlWeb(); |
| 82 |
$html->load('https://slashdot.org/'); |
| 83 |
|
| 84 |
$articles = $html->find('article[data-fhtype="story"]'); |
| 85 |
|
| 86 |
foreach($articles as $article) { |
| 87 |
$item['title'] = $article->find('.story-title', 0)->plaintext; |
| 88 |
$item['intro'] = $article->find('.p', 0)->plaintext; |
| 89 |
$item['details'] = $article->find('.details', 0)->plaintext; |
| 90 |
$items[] = $item; |
| 91 |
} |
| 92 |
|
| 93 |
print_r($items); |
| 94 |
``` |
| 95 |
|
| 96 |
Collects information from [Slashdot](https://slashdot.org/) for further processing. |
| 97 |
|
| 98 |
Note that the combination of CSS selectors and magic methods make the process of parsing HTML documents a simple task that is easy to understand. |