我目前正在聚集一些文本文档。 由于使用了PySpark方法,我正在使用K-means并使用TF-IDF进行数据处理。 现在,我想获取每个群集的前10个字:
当我这样做时:
getTopwords_udf = udf(lambda vector: [ countVectorizerModel.vocabulary[indice] for indice in vector.toArray().tolist().argsort()[-10:][::-1]], ArrayType(StringType()))
predictions.groupBy("prediction").agg(Summarizer.mean(col("features")).alias("means")) \
.withColumn("topWord", getTopwords_udf(col('means'))) \
.select("prediction", "topWord") \
.show(2, truncate=100)
我遇到此错误:
Could not serialize object: Py4JError: An error occurred while calling o225.__getstate__. Trace:
py4j.Py4JException: Method __getstate__([]) does not exist
at py4j.reflection.ReflectionEngine.getMethod(ReflectionEngine.java:318)
at py4j.reflection.ReflectionEngine.getMethod(ReflectionEngine.java:326)
at py4j.Gateway.invoke(Gateway.java:274)
at py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:132)
at py4j.commands.CallCommand.execute(CallCommand.java:79)
at py4j.GatewayConnection.run(GatewayConnection.java:238)
at java.lang.Thread.run(Thread.java:748)
Traceback (most recent call last):
File "/opt/bigpipe/spark/python/lib/pyspark.zip/pyspark/sql/udf.py", line 189, in wrapper
return self(*args)
File "/opt/bigpipe/spark/python/lib/pyspark.zip/pyspark/sql/udf.py", line 167, in __call__
judf = self._judf
File "/opt/bigpipe/spark/python/lib/pyspark.zip/pyspark/sql/udf.py", line 151, in _judf
self._judf_placeholder = self._create_judf()
File "/opt/bigpipe/spark/python/lib/pyspark.zip/pyspark/sql/udf.py", line 160, in _create_judf
wrapped_func = _wrap_function(sc, self.func, self.returnType)
File "/opt/bigpipe/spark/python/lib/pyspark.zip/pyspark/sql/udf.py", line 35, in _wrap_function
pickled_command, broadcast_vars, env, includes = _prepare_for_python_RDD(sc, command)
File "/opt/bigpipe/spark/python/lib/pyspark.zip/pyspark/rdd.py", line 2420, in _prepare_for_python_RDD
pickled_command = ser.dumps(command)
File "/opt/bigpipe/spark/python/lib/pyspark.zip/pyspark/serializers.py", line 597, in dumps
raise pickle.PicklingError(msg)
_pickle.PicklingError: Could not serialize object: Py4JError: An error occurred while calling o225.__getstate__. Trace:
py4j.Py4JException: Method __getstate__([]) does not exist
at py4j.reflection.ReflectionEngine.getMethod(ReflectionEngine.java:318)
at py4j.reflection.ReflectionEngine.getMethod(ReflectionEngine.java:326)
at py4j.Gateway.invoke(Gateway.java:274)
at py4j.commands.AbstractCommand.invokeMethod(AbstractCommand.java:132)
at py4j.commands.CallCommand.execute(CallCommand.java:79)
at py4j.GatewayConnection.run(GatewayConnection.java:238)
at java.lang.Thread.run(Thread.java:748)
我以为是因为类型(从DoubleType到numpy的float类型),所以我也尝试过这样做以查看正在发生的情况
vector_udf = udf(lambda vector: vector.toArray().tolist(), ArrayType(FloatType()))
vector2_udf = udf(lambda vector: vector.sort()[:10], ArrayType(FloatType()))
predictions.groupBy("prediction").agg(Summarizer.mean(col("features")).alias("means")) \
.withColumn("topWord", vector_udf(col('means'))) \
.withColumn("topWord2", vector2_udf(col('topWord'))) \
.select("prediction", "topWord", "topWord2") \
.show(2, truncate=100)
但是我收到此错误TypeError: 'NoneType' object is not subscriptable
答案 0 :(得分:1)
我已经弄清楚了如何使用PySpark从SparseVector到字符串数组获取单词的前X个。 这是我为那些可能感兴趣的人提供的解决方案...
def getTopWordContainer(v):
def getTopWord(vector):
vectorConverted = vector.toArray().tolist()
listSortedDesc= [i[0] for i in sorted(enumerate(vectorConverted), key=lambda x:x[1])][-10:][::-1]
return [v[j] for j in listSortedDesc]
return getTopWord
getTopWordInit = getTopWordContainer(countVectorizerModel.vocabulary)
getTopWord_udf = udf(getTopWordInit, ArrayType(StringType()))
top = predictions.groupBy("prediction").agg(Summarizer.mean(col("features")).alias("means")) \
.withColumn("topWord", getTopWord_udf(col('means'))) \
.select("prediction", "topWord")
我是Spark的初学者,因此,如果您知道如何增强它,请告诉我:)