如何获取Office文档的子类型MIME,而不是在Tika中获取OOXML

时间:2018-07-03 15:05:22

标签: java mime apache-tika

我正在使用Tika验证文件类型,并确保没有人试图以真实文件为幌子发送恶意或伪造文件。为此,我正在使用Apache Tika。但是,即使我将InputStream包装到TikaInputStream中,或者使用OOXMLParser或OfficeParser,它仍然会返回application / x-tika-ooxml而不是application / vnd.openxmlformats-officedocument.wordprocessingml.document。如何访问或获取它以返回子类型?

    public static boolean isValidFileMimeType(TikaInputStream stream, String[] validMimes) {
    Tika tika = new Tika();
    try {
        Metadata meta = new Metadata();
        tika.detect(stream, meta);
        String mimetype = meta.get("Content-Type");
        logger.debug("MIME type from TIKA is : [" + mimetype +"]");
        logger.debug(meta.toString());
        //return isValidFileMimeType(mimetype, validMimes);
        return true;
    } catch (Exception e) {
        logger.error("Error validating InputStream: ", e);
        return false;
    }

public static boolean isValidFileMimeType(MultipartFile file, String[] mimeTypes) {
    TikaInputStream in = null;
    boolean isValidFile = false;
     try {
         in = TikaInputStream.get(file.getInputStream());
        isValidFile = DataValidator.isValidFileMimeType(in, mimeTypes);
    } catch (IOException e) {
        logger.error("Error while validating file mime type: ", e);
    } finally {
        if (in != null) {
            try {
                in.close();
            } catch (IOException e2) {
                logger.error("Error while closing InputStream: ", e2);
            }
        }
    }
        return isValidFile;
}

1 个答案:

答案 0 :(得分:0)

只需导入/使用 Tika 解析器