使用Simple HTML DOM Parser在html标签内获取数据:

时间:2013-08-29 11:58:29

标签: php html-table html-parsing simple-html-dom

我希望获取html标记内的所有信息并将其显示在表格中。我正在使用Simple HTML DOM Parser。我尝试了以下代码,但我只获得了最后一列(列:总计)。如何从其他列获取数据?

foreach($html->find('tr[class="tblRowShade"]') as $div) {
    $key = '';
    $val = '';

    foreach($div->find('*') as $node) {
        if ($node->tag=='td'){
            $key = $node->plaintext;
        }
    }

    $ret[$key] = $val;
}

这是我的表格代码

 <tr class="tblRowShade">
      <td width="12%"><strong>Project</strong></td>
      <td width="38%">&nbsp;</td>
      <td width="25%"><strong>Recipient</strong></td>
      <td width="14%"><strong>Municipality/City</strong></td>
      <td width="11%" nowrap="nowrap" class="td_right"><strong>Implementing Unit</strong></td>
      <td width="11%" nowrap="nowrap" class="td_right"><strong>Release Date</strong></td>
      <td align="right" width="11%" class="td_right"><strong>Total</strong></td>
 </tr>

<tr class="tblRowShade">
      <td colspan="2" >Livelihood Programs</td>
      <td >Basic Espresso and Latte</td>
      <td nowrap="nowrap"></td>
      <td >DOLE - TESDA Regional Office IV-A</td>
      <td nowrap="nowrap">2013-06-11</td>
      <td align="right" nowrap="nowrap" class="td_right">1,500,000</td>
</tr>

2 个答案:

答案 0 :(得分:0)

为什么你有$div->find('*')?你可以试试$div->find('td')。这应该产生正确的结果。否则,您也可以尝试迭代子项:foreach($div->children as $node)

假设您尝试将第一行用作$ key而其余用于数据,您可能只想在第一行中添加th来更改HTML代码,这是您的标题:{{ 1}}。这样您就可以通过<tr><th>…</th></tr>获取密钥。我想使用第一行也没关系。

答案 1 :(得分:0)

正如alamin.ahmed所说,最好搜索td而不是......

这是一个有效的例子:

$text = ' <tr class="tblRowShade">
      <td width="12%"><strong>Project</strong></td>
      <td width="38%">&nbsp;</td>
      <td width="25%"><strong>Recipient</strong></td>
      <td width="14%"><strong>Municipality/City</strong></td>
      <td width="11%" nowrap="nowrap" class="td_right"><strong>Implementing Unit</strong></td>
      <td width="11%" nowrap="nowrap" class="td_right"><strong>Release Date</strong></td>
      <td align="right" width="11%" class="td_right"><strong>Total</strong></td>
 </tr>

<tr class="tblRowShade">
      <td colspan="2" >Livelihood Programs</td>
      <td >Basic Espresso and Latte</td>
      <td nowrap="nowrap"></td>
      <td >DOLE - TESDA Regional Office IV-A</td>
      <td nowrap="nowrap">2013-06-11</td>
      <td align="right" nowrap="nowrap" class="td_right">1,500,000</td>
</tr>';

echo  "<div>Original Text: <xmp>$text</xmp></div>";


//Create a DOM object
$html = new simple_html_dom();
// Load HTML from a string
$html->load($text);


// Find all elements
$rows = $html->find('tr[class="tblRowShade"]');


// Find succeeded
if ($rows) {

    echo count($rows) . " \$rows found !<br />";

    foreach ($rows as $key => $row) {

        echo "<hr />";

        $columns = $row->find('td');

        // Find succeeded
        if ($rows) {

            echo count($columns) . " \$columns found  in \$rows[$key]!<br />";

            foreach ($columns as $col) {

                    echo $col->plaintext . " | ";
                }
        }
        else
            echo " /!\ Find() \$columns failed /!\ ";
    }
}
else
    echo " /!\ Find() \$rows failed /!\ ";

这是上面代码的输出:

enter image description here

您必须知道这两行不包含相同数量的列...然后您必须在程序中处理它。