Question

如果我写这段代码，我会把它作为输出 - ＆gt;第一个：ï»¿ 然后是其他行

try {
    BufferedReader br = new BufferedReader(new FileReader(
            "myFile.txt"));

    String line;
    while (line = br.readLine() != null) {
        System.out.println(line);
    }
    br.close();

} catch (FileNotFoundException e) {
    e.printStackTrace();
} catch (IOException e) {
    e.printStackTrace();
}

我该如何避免呢？

Answer 1

你在第一行得到了字符ï»¿因为这个序列是UTF-8 byte order mark (BOM)。如果文本文件以BOM开头，则很可能是由记事本等Windows程序生成的。

要解决您的问题，我们选择将文件显式读取为UTF-8，而不是默认的系统字符编码（US-ASCII等）：

BufferedReader in = new BufferedReader(
    new InputStreamReader(
        new FileInputStream("myFile.txt"),
        "UTF-8"));

然后在UTF-8中，字节序列ï»¿解码为一个字符，即U + FEFF。此字符是可选的 - 合法的UTF-8文件可能也可能不以它开头。因此，只有当它是U + FEFF时才会跳过第一个字符：

in.mark(1);
if (in.read() != 0xFEFF)
  in.reset();

现在，您可以继续使用其余代码。

Answer 2

问题可能在于使用的编码。试试这个：

BufferedReader in = new BufferedReader(new InputStreamReader(
      new FileInputStream("yourfile"), "UTF-8"));

从文本文件中读取的第一个字符：ï»¿

2 个答案: