无法读取doc URL:无法读取整个标头;读取6个字节;预期的32个字节

时间:2019-05-11 13:12:15

标签: java apache-poi doc

我正在尝试使用POI 3.6版从Web URL读取Word文档。无效的代码:

String url = "http://prevention.cancer.gov/sites/default/files/uploads/clinical_trial/Master-DMP-Template.doc";
InputStream inputStream = new URL(urlString).openStream();
HWPFDocument doc = new HWPFDocument(inputStream);
WordExtractor extractor = new WordExtractor(doc);
String text = extractor.getText();

以上代码导致java.io.IOException:无法读取整个标头;读取6个字节;预期的32个字节

尝试2:有趣的部分是下载文件(只需将URL粘贴在浏览器地址栏中),然后执行类似的代码以在本地读取文档即可:

InputStream inputStream = new FileInputStream("C:\\Users\\me\\Downloads\\Master-DMP-Template (2).doc");
HWPFDocument doc = new HWPFDocument(inputStream);
WordExtractor extractor = new WordExtractor(doc);
System.out.println(extractor.getText());

尝试3:现在是最奇怪的部分。我认为该文件需要首先下载。因此,我首先使用Java下载了该文件,然后执行了前面的代码以本地读取该文档。像第一种情况一样失败!

final String url = "http://prevention.cancer.gov/sites/default/files/uploads/clinical_trial/Master-DMP-Template.doc";
String localPath  = FileUtils.downloadFile("C:\\Users\\me\\Downloads", url);
InputStream inputStream = new FileInputStream(localPath);
HWPFDocument doc = new HWPFDocument(inputStream);
WordExtractor extractor = new WordExtractor(doc);
System.out.println(extractor.getText());

public static String downloadFile(String targetDir, String sourceUrl) throws IOException {
    sourceUrl = StringUtils.removeEnd(sourceUrl, "/");
    String fileName = sourceUrl.substring(sourceUrl.lastIndexOf("/") + 1);
    String targetPath = targetDir + FileUtils.SEPARATOR + fileName;
    InputStream in = new URL(sourceUrl).openStream();
    Files.copy(in, Paths.get(targetPath), StandardCopyOption.REPLACE_EXISTING);
    System.out.println("Downloaded " + sourceUrl + " to " + targetPath);
    return targetPath;
}

有什么想法吗?

一个更新:我创建了一个单独的项目来尝试使用POI 4.1.0。相同的代码(第一次尝试)会导致org.apache.poi.EmptyFileException:提供的文件为空(零字节长)

在按F12键并观察“网络”选项卡后,我尝试将URL粘贴到浏览器中。出现的消息是: 资源被解释为文档,但是以MIME类型application / msword传输:“ https://prevention.cancer.gov/sites/default/files/uploads/clinical_trial/Master-DMP-Template.doc”。

我仍然被困住...

一个更新:正如https://stackoverflow.com/users/3915431/axel-richter所指出的,有一个到https://prevention.cancer.gov/sites/default/files/uploads/clinical_trial/Master-DMP-Template.doc的301重定向。但是,现在我遇到了与Word不相关的奇怪问题。 Followig代码失败:

public static void main(String[] args) {
    try {
        if (args.length > 0 && args[0].equals("disableCertValidation")) {
            SSLUtil.disableCertificateValidation(); // redirect is https
        }
        final String stringURL = "https://prevention.cancer.gov/sites/default/files/uploads/clinical_trial/Master-DMP-Template.doc";
        URL url = new URL(stringURL);
        HttpURLConnection con = (HttpURLConnection) url.openConnection();
        int responseCode = con.getResponseCode();
        System.out.println("Response code: " + responseCode); //301 Moved Permanently
        InputStream in = con.getInputStream();
        HWPFDocument doc = new HWPFDocument(in);
        WordExtractor extractor = new WordExtractor(doc);
        String text = extractor.getText();
        System.out.println(text);
        in.close();
    } catch (IOException e) {
        e.printStackTrace();
    }
}

在不带参数的情况下运行main时,行

int responseCode = con.getResponseCode();

失败,但以下情况除外: javax.net.ssl.SSLHandshakeException:sun.security.validator.ValidatorException:PKIX路径构建失败:sun.security.provider.certpath.SunCertPathBuilderException:无法找到请求的目标的有效证书路径

运行带有disableCertificateValidation参数的代码时,响应代码为404,并且出现以下异常:

java.io.FileNotFoundException:https://prevention.cancer.gov/sites/default/files/uploads/clinical_trial/Master-DMP-Template.doc         在sun.reflect.NativeConstructorAccessorImpl.newInstance0(本机方法)处         在sun.reflect.NativeConstructorAccessorImpl.newInstance(NativeConstructorAccessorImpl.java:62)         在sun.reflect.DelegatingConstructorAccessorImpl.newInstance(DelegatingConstructorAccessorImpl.java:45)         在java.lang.reflect.Constructor.newInstance(Constructor.java:422)         在sun.net.www.protocol.http.HttpURLConnection $ 10.run(HttpURLConnection.java:1890)         在sun.net.www.protocol.http.HttpURLConnection $ 10.run(HttpURLConnection.java:1885)         在java.security.AccessController.doPrivileged(本机方法)         在sun.net.www.protocol.http.HttpURLConnection.getChainedException(HttpURLConnection.java:1884)         在sun.net.www.protocol.http.HttpURLConnection.getInputStream0(HttpURLConnection.java:1457)         在sun.net.www.protocol.http.HttpURLConnection.getInputStream(HttpURLConnection.java:1441)         在sun.net.www.protocol.https.HttpsURLConnectionImpl.getInputStream(HttpsURLConnectionImpl.java:254)         在com.keywords.control.util.TestHTMLParser.main(TestHTMLParser.java:472) 原因:java.io.FileNotFoundException:https://prevention.cancer.gov/sites/default/files/uploads/clinical_trial/Master-DMP-Template.doc         在sun.net.www.protocol.http.HttpURLConnection.getInputStream0(HttpURLConnection.java:1836)         在sun.net.www.protocol.http.HttpURLConnection.getInputStream(HttpURLConnection.java:1441)         在java.net.HttpURLConnection.getResponseCode(HttpURLConnection.java:480)         在sun.net.www.protocol.https.HttpsURLConnectionImpl.getResponseCode(HttpsURLConnectionImpl.java:338)         在com.keywords.control.util.TestHTMLParser.main(TestHTMLParser.java:470)

有什么想法吗?

2 个答案:

答案 0 :(得分:1)

对您的HTTP的初始URL请求导致重定向301 Moved Permanently。因此,我们需要处理并读取新位置。

完整示例:

import java.io.InputStream;
import java.net.URL;
import java.net.HttpURLConnection;

import org.apache.poi.hwpf.HWPFDocument;
import org.apache.poi.hwpf.extractor.WordExtractor;

public class OpenHWPFFromURL {

 public static void main(String[] args) throws Exception {

  String stringURL = "http://prevention.cancer.gov/sites/default/files/uploads/clinical_trial/Master-DMP-Template.doc";

  URL url = new URL(stringURL);
  HttpURLConnection con = (HttpURLConnection)url.openConnection();

  int responseCode = con.getResponseCode();
  System.out.println(responseCode); //301 Moved Permanently

  if (responseCode != HttpURLConnection.HTTP_OK) {
   if (responseCode == HttpURLConnection.HTTP_MOVED_TEMP
       || responseCode == HttpURLConnection.HTTP_MOVED_PERM
       || responseCode == HttpURLConnection.HTTP_SEE_OTHER) {
    url = new URL(con.getHeaderField("Location")); //get new location
    con = (HttpURLConnection)url.openConnection();
   }   
  }

  InputStream in = con.getInputStream();
  HWPFDocument doc = new HWPFDocument(in);
  WordExtractor extractor = new WordExtractor(doc);
  String text = extractor.getText();

  System.out.println(text);

 }
}

注意:如果重定向也将协议(从HttpURLConnection.setFollowRedirects更改为true,则仅将HTTP设置为HTTPS(也是默认设置)将无济于事。例)。确实在这里也是这种情况。因此,我们需要手动获取新位置,如我的代码所示。

答案 1 :(得分:0)

此代码new URL(urlString).openStream()返回InputStream look here,而不是FileInputStream,如下所示:

InputStream inputStream = new FileInputStream("C:\\Users\\me\\Downloads\\Master...")

也许在这种区别上有问题吗?