Postgres - 按月顺序进行队列分析,而不是在任何后期月份进行

时间:2017-12-05 16:40:42

标签: sql postgresql cluster-analysis analytics

我正在进行同期群分析,可以让用户群进行检查,然后查看他们是否会在接下来的几个月内进行交易。但我想要这样:

12月份的那个小组,谁在1月份进行了交易; 12月交易的Jan集团,他在2月份进行交易。基本上我跟踪客户群的衰退

我不想要的是那些在12月之后的任何月份返回的,这是:

WITH start_sample AS (
SELECT
  user_fk,
  created_at AS start_sample_date
  FROM transactions
    WHERE created_at >= '2016-11-01' AND created_at < '2016-12-01'
      GROUP BY user_fk,
        start_sample_date),

start_sample_min AS (
SELECT
  user_fk,
  MIN(start_sample_date) AS first_transaction
    FROM start_sample
      GROUP BY user_fk
  )

SELECT
  DATE_TRUNC('month', created_at) AS transacting_month,
  COUNT(DISTINCT user_fk)
    FROM transactions
        WHERE created_at >= '2016-11-01'
        AND t.user_fk IN(SELECT user_fk FROM start_sample_min)
          GROUP BY transacting_month
            ORDER BY transacting_month;

然后我制作了一个流失模型,看它是否能得到我需要的东西,但它没有:

WITH monthly_users AS (
    SELECT
      user_fk AS monthly_user_fk,
      DATE_TRUNC('month', created_at) AS month
        FROM transactions
          WHERE created_at >= '2016-11-01' AND created_at < '2017-12-01'
            GROUP BY monthly_user_fk, month
            ORDER BY monthly_user_fk, month
),

lag_lead AS (
  SELECT
    monthly_user_fk,
    month,
    LAG(month,1) OVER (PARTITION BY monthly_user_fk ORDER BY month) AS lag,
    LEAD(month,1) OVER (PARTITION BY monthly_user_fk ORDER BY month) AS lead
      FROM monthly_users),

lag_lead_with_diffs AS (
  SELECT
    monthly_user_fk,
    month,
    lag AS previous_month,
    lead AS next_month,
    EXTRACT(EPOCH FROM (month - lag)/86400)::INT AS lag_size,
    EXTRACT(EPOCH FROM (lead - month)/86400)::INT AS lead_size
      FROM lag_lead
  ),

calculated AS (
      SELECT
      month,
      CASE WHEN previous_month IS NULL THEN 'ACTIVATION'
          WHEN lag_size <= 31 THEN 'ACTIVE'
          WHEN lag_size > 31 THEN 'RETURN' END AS this_month_values,
      CASE WHEN (lead_size > 31 OR lead_size IS NULL) THEN 'CHURN' ELSE NULL END AS next_month_churn,
      COUNT(DISTINCT monthly_user_fk) AS c_d_users
   FROM lag_lead_with_diffs
  GROUP BY month, 2, 3
)

SELECT
  month,
  this_month_values,
  SUM(c_d_users) AS distinct_users
  FROM calculated
  GROUP BY month, this_month_values
UNION
SELECT month + INTERVAL '1 month',
  'CHURN',
  SUM(c_d_users)
  FROM calculated
    WHERE next_month_churn IS NOT NULL
      GROUP BY month + INTERVAL '1 month', 2
        HAVING (EXTRACT(EPOCH FROM (month + INTERVAL '1 month'))) < 1512086400
          ORDER BY month, this_month_values;

然而,初始组并未解决此问题。活动组每月滚动一次。

据我所知,上述内容可能比我要求的更复杂,但我似乎无法理解它

提前致谢

1 个答案:

答案 0 :(得分:2)

也许这就是你要找的东西:

with Monthly_Users as (
select user_fk
     , date_trunc('month',created_at) as month
     , (date_part('year', created_at) - 2016) * 12
     + date_part('month', created_at) - 11 as Months_Between
  from transactions
 where created_at between date '2016-11-01'
                      and date '2017-12-01'
 group by user_fk, month, months_between
), t2 as (
select Monthly_Users.*
     , count(*) over (partition by user_fk
                          order by month rows between unbounded preceding
                                                  and 1 preceding) prev_rec_cnt
  from Monthly_Users
)
select month
     , count(*)
  from t2
 where Months_Between = Prev_Rec_Cnt
 group by month
 order by month;

在此查询中,Monthly_Users CTE与您的一样,但会计算Months_Between created_at日期和初始开始日期的数量。在第二个公用表表达式中,我计算当前month s记录之前每个user_fk的出现次数。最后,在输出查询中,我将结果限制为Months_Between值与Prev_Rec_Cnt值匹配的记录。任何错过的月份都会导致Prev_Rec_Cnt值与Months_Between值不匹配,因此您每个月都可以看到user_fk值的下降。