猪 - 在其他'中获得最佳和团体休息。

时间:2015-08-03 19:34:22

标签: hadoop apache-pig hdfs

我有分组和汇总的数据,看起来像这样 -

Date Country Browser Count
---- ------- ------- -----
2015-07-11,US,Chrome,13
2015-07-11,US,Opera Mini,1
2015-07-11,US,Firefox,2
2015-07-11,US,IE,1
2015-07-11,US,Safari,1
...
2015-07-11,UK,Chrome Mobile,1026
2015-07-11,UK,IE,455
2015-07-11,UK,Mobile Safari,4782
2015-07-11,UK,Mobile Firefox,40
...
2015-07-11,DE,Android browser,1316
2015-07-11,DE,Opera Mini,3
2015-07-11,DE,PS4 Web browser,11

我希望每个国家/地区获得前n个浏览器(按计数),并希望将其余浏览器汇总到其他'其他'。我查看了Pig的内置TOP功能,但是如何在其他地方进行分组。我想要的结果,例如(n = 2) - >

2015-07-11,US,Chrome,13
2015-07-11,US,Firefox,2
2015-07-11,US,Other,3

最好的方法是什么?

1 个答案:

答案 0 :(得分:1)

好的..这个要求很好..

我只是在Pig脚本的LOAD语句中使用你的输入。

输入:

2015-07-11,US,Chrome,13
2015-07-11,US,Opera Mini,1
2015-07-11,US,Firefox,2
2015-07-11,US,IE,1
2015-07-11,US,Safari,1
2015-07-11,UK,Chrome Mobile,1026
2015-07-11,UK,IE,455
2015-07-11,UK,Mobile Safari,4782
2015-07-11,UK,Mobile Firefox,40
2015-07-11,DE,Android browser,1316
2015-07-11,DE,Opera Mini,3
2015-07-11,DE,PS4 Web browser,11
2015-07-11,US,Chrome,13
2015-07-11,US,Firefox,2
2015-07-11,US,Other,3

以下是对此的编码。

你可以将n参数的值传递给pig脚本,目前我在LIMIT语句中为n设置值2(即n = 2)。

实际上我在下面的代码中硬编码了n = 2.

records     = LOAD '/user/cloudera/inputfiles/entries.txt' USING PigStorage(',') as (dt:chararray,country:chararray,browser:chararray,count:int);

records_each    = FOREACH(GROUP records BY (dt,country,browser)) GENERATE flatten(group) AS (dt,country,browser), MAX(records.count) as counts;

records_grp_order = ORDER records_each BY dt ASC , country  ASC , counts DESC;

records_grp     = GROUP records_grp_order BY (dt, country);

rec_each    = FOREACH records_grp {

               top_2_recs = LIMIT records_grp_order  2;
               generate  MAX(top_2_recs.dt) AS temp_dt, MAX(top_2_recs.country) AS temp_country, flatten(top_2_recs.browser) AS temp_browser;

            };
rec_join    =  JOIN records_each BY (dt,country,browser)  left outer , rec_each BY (temp_dt,temp_country,temp_browser);

rec_join_each   = FOREACH rec_join generate dt,country, (temp_browser is not null ? browser : 'OTHERS') AS browser, counts AS counts;

rec_final_grp   = GROUP rec_join_each BY (dt,country,browser);

final_output    = FOREACH rec_final_grp generate flatten(group) AS (dt,country,browser), SUM(rec_join_each.counts) AS total_counts;

sorted_output   = ORDER final_output BY  dt ASC , country  ASC, total_counts DESC;

dump sorted_output;

输出

(2015-07-11,DE,Android browser,1316)
(2015-07-11,DE,PS4 Web browser,11)
(2015-07-11,DE,OTHERS,3)
(2015-07-11,UK,Mobile Safari,4782)
(2015-07-11,UK,Chrome Mobile,1026)
(2015-07-11,UK,OTHERS,495)
(2015-07-11,US,Chrome,13)
(2015-07-11,US,OTHERS,3)
(2015-07-11,US,Firefox,2)