Spark-将JSON数组对象转换为字符串数组

时间:2019-07-14 21:04:24

标签: apache-spark dataframe pyspark apache-spark-sql

作为我的数据框的一部分,其中一列具有以下方式的数据

  

[{“ text”:“ Tea”},{“ text”:“ GoldenGlobes”}]

我想将其转换为只是字符串数组。

  

[“ Tea”,“ GoldenGlobes”]

请让我知道,该怎么做?

3 个答案:

答案 0 :(得分:0)

如果您的列的类型是数组,则类似这样的方法应该起作用(未经测试):

from pyspark.sql import functions as F
from pyspark.sql import types as T

c = F.array([F.get_json_object(F.col("colname")[0], '$.text')),  
             F.get_json_object(F.col("colname")[1], '$.text'))])

df = df.withColumn("new_col", c)

或者如果长度不是固定的(没有udf,我看不到解决方案):

F.udf(T.ArrayType())
def get_list(x):
    o_list = []
    for elt in x:
        o_list.append(elt["text"])
    return o_list

df = df.withColumn("new_col", get_list("colname"))

答案 1 :(得分:0)

请参见下面的示例,其中不包含udf

import pyspark.sql.functions as f
from pyspark import Row
from pyspark.shell import spark
from pyspark.sql.types import ArrayType, StructType, StructField, StringType

df = spark.createDataFrame([
    Row(values='[{"text":"Tea"},{"text":"GoldenGlobes"}]'),
    Row(values='[{"text":"GoldenGlobes"}]')
])

schema = ArrayType(StructType([
    StructField('text', StringType())
]))

df \
    .withColumn('array_of_str', f.from_json(f.col('values'), schema).text) \
    .show()

输出:

+--------------------+-------------------+
|              values|       array_of_str|
+--------------------+-------------------+
|[{"text":"Tea"},{...|[Tea, GoldenGlobes]|
|[{"text":"GoldenG...|     [GoldenGlobes]|
+--------------------+-------------------+

答案 2 :(得分:0)

共享Java语法:

import static org.apache.spark.sql.functions.from_json;
import static org.apache.spark.sql.functions.get_json_object;
import static org.apache.spark.sql.functions.col;
import org.apache.spark.sql.types.StructType;
import org.apache.spark.sql.types.DataTypes;
import org.apache.spark.sql.types.StructField;
import static org.apache.spark.sql.types.DataTypes.StringType;

Dataset<Row> df = getYourDf();

StructType structschema =
                DataTypes.createStructType(
                        new StructField[] {
                                DataTypes.createStructField("text", StringType, true)
                        });

ArrayType schema = new ArrayType(structschema,true);


df = df.withColumn("array_of_str",from_json(col("colname"), schema).getField("text"));