Skip to content

[Python] Table.from_pandas creates duplicate column names if the dataframe already contains __index_level_i__ columns #46179

Description

@jorisvandenbossche

Describe the bug, including details regarding any error messages, version, and platform.

The pandas -> arrow conversion adds a __inex_level_i__ column if the dataframe has an unnamed it wants to preserve (i.e. if it is not just a pandas RangeIndex). But if your dataframe already has such a column, you end up with a duplicate field:

In [40]: df = pd.DataFrame({"col": [1, 2, 3], "__index_level_0__": [1, 2, 3]}, index=[2, 3, 4])

In [41]: df
Out[41]: 
   col  __index_level_0__
2    1                  1
3    2                  2
4    3                  3

In [42]: pa.table(df)
Out[42]: 
pyarrow.Table
col: int64
__index_level_0__: int64
__index_level_0__: int64
----
col: [[1,2,3]]
__index_level_0__: [[1,2,3]]
__index_level_0__: [[2,3,4]]

We could have it bump the integer number in the generated column? (although we would have to check how that works in the full roundtrip then)

Component(s)

Python

Related issues

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions