带有DOMXPath的PHP - 如何从这个html树中选择和计数

时间:2016-01-15 18:40:42

标签: php domxpath

我需要计算这些项目中有多少是开放的,有四种类型:简单,中等,困难和不想要。所有这些类型都是div中的值。我需要排除' Not-Wanted'来自伯爵的类型。注意'打开'并且'关闭'值周围有不同的空格数。这是html结构:

<table>
    <tbody>
        <tr>
            <td>
                <div>Difficult</div>
            </td>
            <td>Name</td>
            <td>  Open </td>
        </tr>
        <tr>
            <td>
                <div>Easy</div>
            </td>
            <td>Name</td>
            <td> Closed  </td>
        </tr>
        <tr>
            <td>
                <div>Easy</div>
            </td>
            <td>Name</td>
            <td>   Open   </td>
        </tr>
        <tr>
            <td>
                <div>Medium</div>
            </td>
            <td>Name</td>
            <td>Open </td>
        </tr>
        <tr>
            <td>
                <div>Easy</div>
            </td>
            <td>Name</td>
            <td> Open     </td>
        </tr>
        <tr>
            <td>
                <div>Medium</div>
            </td>
            <td>Name</td>
            <td>  Closed</td>
        </tr>
        <tr>
            <td>
                <div>Easy</div>
            </td>
            <td>Name</td>
            <td>Closed </td>
        </tr>
        <tr>
            <td>
                <div>Not-wanted</div>
            </td>
            <td>Name</td>
            <td> Open </td>
        </tr>
        <tr>
            <td>
                <div>Difficult</div>
            </td>
            <td>Name</td>
            <td>Open</td>
        </tr>
        ............

这是我尝试解决问题的方法之一。这显然是错的,但我不知道如何做对。

$doc = new DOMDocument();
$doc->loadHtmlFile('http://www.nameofsite.com');
$doc->preserveWhiteSpace = false;
$xpath = new DOMXPath($doc);

$elements = $xpath->query("/html/body/div[1]/div/section/div/section/article/div/div[1]/div/div/div[2]/div[1]/div[2]/div/section/div/div/table/tbody/tr");

$count = 0;
foreach ($elements as $element) {
    if ($element->childNodes->nodeValue != 'Not-wanted') {
        if ($element->childNodes->nodeValue === 'open') {
            $count++;
        }
    }
}

echo $count;

我对DOMXPath有一个非常基本的知识,所以它对我来说太复杂了,因为我只能创建简单的查询。

有人可以帮忙吗?

提前致谢。

1 个答案:

答案 0 :(得分:1)

根据您示例中的数据,我认为您可以将xpath表达式调整为此值,以获得符合您条件的所有<tr>

  

// table / tbody / tr [normalize-space(td [3] / text())=&#39;打开&#39;和   td [1] / div / text()!=&#39;不想要&#39;]

$elements的类型为DOMNodeList,然后您可以获取length属性以获取列表中的节点数。

例如:

$source = <<<SOURCE
<table>
    <tbody>
        <tr>
            <td>
                <div>Difficult</div>
            </td>
            <td>Name</td>
            <td>  Open </td>
        </tr>
        <tr>
            <td>
                <div>Easy</div>
            </td>
            <td>Name</td>
            <td> Closed  </td>
        </tr>
        <tr>
            <td>
                <div>Easy</div>
            </td>
            <td>Name</td>
            <td>   Open   </td>
        </tr>
        <tr>
            <td>
                <div>Medium</div>
            </td>
            <td>Name</td>
            <td>Open </td>
        </tr>
        <tr>
            <td>
                <div>Easy</div>
            </td>
            <td>Name</td>
            <td> Open     </td>
        </tr>
        <tr>
            <td>
                <div>Medium</div>
            </td>
            <td>Name</td>
            <td>  Closed</td>
        </tr>
        <tr>
            <td>
                <div>Easy</div>
            </td>
            <td>Name</td>
            <td>Closed </td>
        </tr>
        <tr>
            <td>
                <div>Not-wanted</div>
            </td>
            <td>Name</td>
            <td> Open </td>
        </tr>
        <tr>
            <td>
                <div>Difficult</div>
            </td>
            <td>Name</td>
            <td>Open</td>
        </tr>
    </tbody>
</table>
SOURCE;

$doc = new DOMDocument();
$doc->loadHTML($source);
$doc->preserveWhiteSpace = false;
$xpath = new DOMXPath($doc);
$elements = $xpath->query("//table/tbody/tr[normalize-space(td[3]/text()) = 'Open' and td[1]/div/text() != 'Not-wanted']");
echo $elements->length;

这将导致:

  

5

Demo