使用另一个数据框在数据框中创建子列

时间:2019-11-22 09:21:16

标签: python python-3.x pandas numpy

我是python和pandas的新手。在这里,我有一个以下数据框。

did           features   offset   word   JAPE_feature  manual_feature 
0             200         0        aa      200          200 
0             200         11       bf      200          200
0             200         12       vf      100          100
0             100         13       rw      2200         2200
0             100         14       asd     2600         100 
0             2200        16       dsdd    2200         2200
0             2600        18       wd      2200         2600 
0             2600        20       wsw     2600         2600 
0             4600        21        sd     4600         4600

现在,我有一个数组,其中包含可以为该ID显示的所有特征值。

feat = [100,200,2200,2600,156,162,4600,100]

现在,我正在尝试创建一个看起来像这样的数据框,

id                    Features 
           100   200   2200   2600  156   162    4600  100
0           0     1      0     0     0     0      0     0
1           0     1      0     0     0     0      0     0
2           0     1      0     0     0     0      0     0
3           0     1      0     0     0     0      0     0
4           1     0      0     0     0     0      0     0
5           1     0      0     0     0     0      0     0
7           0     0      1     0     0     0      0     0
8           0     0      0     1     0     0      0     0
9           0     0      0     1     0     0      0     0
10          0     0      0     0     0     0      1     0

所以,在进行比较时,

feature_manual
     1 
     1  
     0 
     0
     1
     1
     1
     1
     1

Here compairing the features and the manual_feature columns. if values are same then 1 or else 0. so 200 and 200 for 0 is same in both so 1 

因此,这是预期的输出。在这里,我正在尝试在新的csv中为该功能添加值1,并为其他0添加值。

So, it is by row by row.

因此,如果我们在第一行中检查该特征为200,则200处为1,其他为0。

有人可以帮助我吗?

我尝试的是

mux = pd.MultiIndex.from_product([['features'],feat)
df = pd.DataFrame(data, columns=mux)

SO,此处创建子列,但删除所有其他值。有人可以帮我吗?

3 个答案:

答案 0 :(得分:2)

get_dummiesDataFrame.reindex一起使用:

feat = [100,200,2200,2600,156,162,4600,100]
df = df.join(pd.get_dummies(df.pop('features')).reindex(feat, axis=1, fill_value=0))
print (df)
   id  100  200  2200  2600  156  162  4600  100
0   0    0    1     0     0    0    0     0    0
1   1    0    1     0     0    0    0     0    0
2   2    0    1     0     0    0    0     0    0
3   4    1    0     0     0    0    0     0    1
4   5    1    0     0     0    0    0     0    1
5   7    0    0     1     0    0    0     0    0
6   8    0    0     0     1    0    0     0    0
7   9    0    0     0     1    0    0     0    0
8  10    0    0     0     0    0    0     1    0

如果需要MultiIndex,则仅将mux传递给reindex,还将id列转换为index

feat = [100,200,2200,2600,156,162,4600,100]
mux = pd.MultiIndex.from_product([['features'],feat])

df = pd.get_dummies(df.set_index('id')['features']).reindex(mux, axis=1, fill_value=0)
print (df)
   features                                   
       100  200  2200 2600 156  162  4600 100 
id                                            
0         0    0    0    0    0    0    0    0
1         0    0    0    0    0    0    0    0
2         0    0    0    0    0    0    0    0
4         0    0    0    0    0    0    0    0
5         0    0    0    0    0    0    0    0
7         0    0    0    0    0    0    0    0
8         0    0    0    0    0    0    0    0
9         0    0    0    0    0    0    0    0
10        0    0    0    0    0    0    0    0

编辑:

cols = ['features', 'JAPE_feature', 'manual_feature']

df = pd.get_dummies(df, columns=cols)
df.columns = df.columns.str.rsplit('_',1, expand=True)
print (df)
  did offset  word features                    JAPE_feature                \
  NaN    NaN   NaN      100 200 2200 2600 4600          100 200 2200 2600   
0   0      0    aa        0   1    0    0    0            0   1    0    0   
1   0     11    bf        0   1    0    0    0            0   1    0    0   
2   0     12    vf        0   1    0    0    0            1   0    0    0   
3   0     13    rw        1   0    0    0    0            0   0    1    0   
4   0     14   asd        1   0    0    0    0            0   0    0    1   
5   0     16  dsdd        0   0    1    0    0            0   0    1    0   
6   0     18    wd        0   0    0    1    0            0   0    1    0   
7   0     20   wsw        0   0    0    1    0            0   0    0    1   
8   0     21    sd        0   0    0    0    1            0   0    0    0   

       manual_feature                     
  4600            100 200 2200 2600 4600  
0    0              0   1    0    0    0  
1    0              0   1    0    0    0  
2    0              1   0    0    0    0  
3    0              0   0    1    0    0  
4    0              1   0    0    0    0  
5    0              0   0    1    0    0  
6    0              0   0    0    1    0  
7    0              0   0    0    1    0  
8    1              0   0    0    0    1  

如果要避免没有MultIndex的列的列中MultiIndex的值丢失:

cols = ['features', 'JAPE_feature', 'manual_feature']
df = df.set_index(df.columns.difference(cols).tolist())

df = pd.get_dummies(df, columns=cols)
df.columns = df.columns.str.rsplit('_',1, expand=True)
print (df)
                features                    JAPE_feature                     \
                     100 200 2200 2600 4600          100 200 2200 2600 4600   
did offset word                                                               
0   0      aa          0   1    0    0    0            0   1    0    0    0   
    11     bf          0   1    0    0    0            0   1    0    0    0   
    12     vf          0   1    0    0    0            1   0    0    0    0   
    13     rw          1   0    0    0    0            0   0    1    0    0   
    14     asd         1   0    0    0    0            0   0    0    1    0   
    16     dsdd        0   0    1    0    0            0   0    1    0    0   
    18     wd          0   0    0    1    0            0   0    1    0    0   
    20     wsw         0   0    0    1    0            0   0    0    1    0   
    21     sd          0   0    0    0    1            0   0    0    0    1   

                manual_feature                     
                           100 200 2200 2600 4600  
did offset word                                    
0   0      aa                0   1    0    0    0  
    11     bf                0   1    0    0    0  
    12     vf                1   0    0    0    0  
    13     rw                0   0    1    0    0  
    14     asd               1   0    0    0    0  
    16     dsdd              0   0    1    0    0  
    18     wd                0   0    0    1    0  
    20     wsw               0   0    0    1    0  
    21     sd                0   0    0    0    1 

编辑:

如果要按manual_feature列比较列表中的某些列,请使用DataFrame.eq转换为整数:

cols = ['JAPE_feature', 'features']
df1 = df[cols].eq(df['manual_feature'], axis=0).astype(int)
print (df1)
   JAPE_feature  features
0             1         1
1             1         1
2             1         0
3             1         0
4             0         1
5             1         1
6             0         1
7             1         1
8             1         1 

答案 1 :(得分:0)

花哨的解决方案较少,但也许更容易理解:

首先,将将决定您选择哪个功能的功能放在称为list_features的列表中。

然后:

# List all the features possible and create an empty df
feat = [100,200,2200,2600,156,162,4600,100]
df_final= pd.DataFrame({x:[] for x in feat})

# Fill the df little by little
for x in list_features:
    df_final = df_final.append({y:1 if x==y else 0 for y in feat }, ignore_index=True)

答案 2 :(得分:0)

这些类型的问题可以通过多种方式解决。但是在这里,我正在使用简单的方法来解决它。创建具有这些功能列表的df作为列名,并使用一些比较逻辑将df更新为0和1。您可以使用其他逻辑来避免使用for循环。

import pandas as pd

data = {'id':[0,1,2,3,4,5,7,8,9,10],
'features':[200, 200, 200, 200, 100, 100, 2200, 2600, 2600, 4600]}

df1 = pd.DataFrame(data)

features_list = [100,200,2200,2600,156,162,4600]
id_list = df1.id.to_list()

df2 = pd.DataFrame(columns=features_list)
list2 = list()

for i in id_list:
    list1 = list()
    for k in df2.columns:
        if df1[df1.id == i].features.iloc[0] == k:
            list1.append(1)
        else:
            list1.append(0)
    list2.append(list1)

for i in range (0,len(list2)):
    df2.loc[i] = list2[i]

df2.insert(0, "id", id_list)   

>>>(df2)
   id 100 200 2200 2600 156 162 4600
0   0   0   1    0    0   0   0    0
1   1   0   1    0    0   0   0    0
2   2   0   1    0    0   0   0    0
3   3   0   1    0    0   0   0    0
4   4   1   0    0    0   0   0    0
5   5   1   0    0    0   0   0    0
6   7   0   0    1    0   0   0    0
7   8   0   0    0    1   0   0    0
8   9   0   0    0    1   0   0    0
9  10   0   0    0    0   0   0    1