[{"data":1,"prerenderedAt":1791},["ShallowReactive",2],{"doc:\u002Fgetting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002Fpandas-vs-polars-for-excel-workflows":3,"surround:\u002Fgetting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002Fpandas-vs-polars-for-excel-workflows":1783},{"id":4,"title":5,"body":6,"dateModified":1754,"datePublished":1754,"description":1755,"extension":1756,"faq":1757,"meta":1768,"navigation":229,"path":1776,"seo":1777,"slug":1779,"stem":1780,"type":1781,"__hash__":1782},"docs\u002Fgetting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002Fpandas-vs-polars-for-excel-workflows\u002Findex.md","pandas vs Polars for Excel Workflows",{"type":7,"value":8,"toc":1740},"minimark",[9,19,118,123,160,181,185,188,328,347,351,354,433,635,650,654,664,780,800,804,807,878,887,891,909,1073,1088,1092,1095,1294,1313,1317,1430,1434,1534,1537,1544,1645,1652,1656,1659,1663,1670,1676,1682,1688,1694,1698,1736],[10,11,12,13,18],"p",{},"Polars arrived as a faster, stricter alternative to pandas, and for the middle of an Excel pipeline\n— the joins, groupings and reshaping between reading a workbook and writing one — it usually\ndeserves the reputation. At the two edges the picture is more nuanced: both libraries delegate the\nactual spreadsheet parsing to the same handful of engines, so \"which is faster at Excel\" is often a\nquestion about calamine rather than about either library. This guide, part of\n",[14,15,17],"a",{"href":16},"\u002Fgetting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002F","Choosing a Python Excel Library",",\ncompares them where the difference is real.",[20,21,29,30,29,34,29,38,29,45,29,52,29,62,29,68,29,73,29,80,29,85,29,89,29,93,29,96,29,100,29,103,29,106,29,110,29,113],"svg",{"viewBox":22,"role":23,"ariaLabelledBy":24,"xmlns":27,"style":28},"0 0 760 232","img",[25,26],"pvp-edges-t","pvp-edges-d","http:\u002F\u002Fwww.w3.org\u002F2000\u002Fsvg","width:100%;max-width:760px;height:auto;display:block;margin:1.5rem auto;font-family:Inter,ui-sans-serif,system-ui,sans-serif","\n  ",[31,32,33],"title",{"id":25},"Where the two libraries actually differ in an Excel pipeline",[35,36,37],"desc",{"id":26},"Reading and writing are shared ground because both libraries delegate to calamine and xlsxwriter. The transform in the middle is where Polars and pandas genuinely differ.",[39,40],"rect",{"x":41,"y":41,"width":42,"height":43,"fill":44},"0","760","232","#ffffff",[46,47,51],"text",{"x":48,"y":49,"style":50},"380.0","32","font-size:13px;font-weight:600;fill:var(--muted,#5b6780);text-anchor:middle","one pipeline, three stages",[39,53],{"x":54,"y":55,"width":56,"height":57,"rx":58,"fill":59,"stroke":60,"style":61},"24.0","74","208.0","96","12","#e7ebef","var(--line,#cdd5e6)","stroke-width:2px",[46,63,67],{"x":64,"y":65,"style":66},"128.0","114","font-size:14px;font-weight:700;fill:var(--muted,#5b6780);text-anchor:middle","read",[46,69,72],{"x":64,"y":70,"style":71},"136","font-size:11.5px;font-weight:400;fill:var(--muted,#5b6780);text-anchor:middle","both use calamine",[74,75],"line",{"x1":76,"y1":77,"x2":78,"y2":77,"stroke":79,"style":61},"237.0","122.0","269.0","var(--brand,#5b5cf0)",[81,82],"polygon",{"points":83,"fill":84},"269.0,122.0 260.0,117.0 260.0,127.0","#5b5cf0",[39,86],{"x":87,"y":55,"width":56,"height":57,"rx":58,"fill":88,"stroke":79,"style":61},"276.0","#f0f4ff",[46,90,92],{"x":48,"y":65,"style":91},"font-size:14px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:middle","transform",[46,94,95],{"x":48,"y":70,"style":71},"the real difference",[74,97],{"x1":98,"y1":77,"x2":99,"y2":77,"stroke":79,"style":61},"489.0","521.0",[81,101],{"points":102,"fill":84},"521.0,122.0 512.0,117.0 512.0,127.0",[39,104],{"x":105,"y":55,"width":56,"height":57,"rx":58,"fill":59,"stroke":60,"style":61},"528.0",[46,107,109],{"x":108,"y":65,"style":66},"632.0","write",[46,111,112],{"x":108,"y":70,"style":71},"both use xlsxwriter",[46,114,117],{"x":48,"y":115,"style":116},"210","font-size:12.5px;font-weight:400;fill:var(--muted,#5b6780);text-anchor:middle","the edges are shared; the middle is the choice",[119,120,122],"h2",{"id":121},"prerequisites","Prerequisites",[124,125,130],"pre",{"className":126,"code":127,"language":128,"meta":129,"style":129},"language-bash shiki shiki-themes github-light github-dark-high-contrast","pip install polars pandas fastexcel xlsxwriter openpyxl\n","bash","",[131,132,133],"code",{"__ignoreMap":129},[134,135,137,141,145,148,151,154,157],"span",{"class":74,"line":136},1,[134,138,140],{"class":139},"sMTad","pip",[134,142,144],{"class":143},"srMev"," install",[134,146,147],{"class":143}," polars",[134,149,150],{"class":143}," pandas",[134,152,153],{"class":143}," fastexcel",[134,155,156],{"class":143}," xlsxwriter",[134,158,159],{"class":143}," openpyxl\n",[10,161,162,165,166,169,170,173,174,177,178,180],{},[131,163,164],{},"fastexcel"," is the binding Polars uses for Excel reads (it wraps calamine); ",[131,167,168],{},"xlsxwriter"," is what\n",[131,171,172],{},"write_excel"," builds on. Both are optional extras, and ",[131,175,176],{},"read_excel","\u002F",[131,179,172],{}," raise a clear\nimport error if they are missing.",[119,182,184],{"id":183},"reading-the-same-workbook-with-each","Reading the same workbook with each",[10,186,187],{},"The calls are close to identical, and so is the time they take when both use the same parser.",[124,189,193],{"className":190,"code":191,"language":192,"meta":129,"style":129},"language-python shiki shiki-themes github-light github-dark-high-contrast","import pandas as pd\nimport polars as pl\n\npdf = pd.read_excel(\"sales.xlsx\", sheet_name=\"Raw\", engine=\"calamine\")\npdf_types = pdf.dtypes\n\npldf = pl.read_excel(\"sales.xlsx\", sheet_name=\"Raw\")\nprint(pldf.schema)\nprint(pdf_types)\n","python",[131,194,195,211,224,231,271,282,287,310,320],{"__ignoreMap":129},[134,196,197,201,205,208],{"class":74,"line":136},[134,198,200],{"class":199},"s-kum","import",[134,202,204],{"class":203},"skGVy"," pandas ",[134,206,207],{"class":199},"as",[134,209,210],{"class":203}," pd\n",[134,212,214,216,219,221],{"class":74,"line":213},2,[134,215,200],{"class":199},[134,217,218],{"class":203}," polars ",[134,220,207],{"class":199},[134,222,223],{"class":203}," pl\n",[134,225,227],{"class":74,"line":226},3,[134,228,230],{"emptyLinePlaceholder":229},true,"\n",[134,232,234,237,240,243,246,249,253,255,258,260,263,265,268],{"class":74,"line":233},4,[134,235,236],{"class":203},"pdf ",[134,238,239],{"class":199},"=",[134,241,242],{"class":203}," pd.read_excel(",[134,244,245],{"class":143},"\"sales.xlsx\"",[134,247,248],{"class":203},", ",[134,250,252],{"class":251},"sa561","sheet_name",[134,254,239],{"class":199},[134,256,257],{"class":143},"\"Raw\"",[134,259,248],{"class":203},[134,261,262],{"class":251},"engine",[134,264,239],{"class":199},[134,266,267],{"class":143},"\"calamine\"",[134,269,270],{"class":203},")\n",[134,272,274,277,279],{"class":74,"line":273},5,[134,275,276],{"class":203},"pdf_types ",[134,278,239],{"class":199},[134,280,281],{"class":203}," pdf.dtypes\n",[134,283,285],{"class":74,"line":284},6,[134,286,230],{"emptyLinePlaceholder":229},[134,288,290,293,295,298,300,302,304,306,308],{"class":74,"line":289},7,[134,291,292],{"class":203},"pldf ",[134,294,239],{"class":199},[134,296,297],{"class":203}," pl.read_excel(",[134,299,245],{"class":143},[134,301,248],{"class":203},[134,303,252],{"class":251},[134,305,239],{"class":199},[134,307,257],{"class":143},[134,309,270],{"class":203},[134,311,313,317],{"class":74,"line":312},8,[134,314,316],{"class":315},"sP0c6","print",[134,318,319],{"class":203},"(pldf.schema)\n",[134,321,323,325],{"class":74,"line":322},9,[134,324,316],{"class":315},[134,326,327],{"class":203},"(pdf_types)\n",[10,329,330,331,334,335,338,339,342,343,346],{},"The differences show in the schema, not the clock. Polars has a real ",[131,332,333],{},"String"," type rather than\n",[131,336,337],{},"object",", distinguishes ",[131,340,341],{},"Int64"," from ",[131,344,345],{},"Float64"," without silently promoting on a missing value, and\nhas no index — so nothing is quietly carried along that you did not ask for. On an Excel file that\nlast point matters more than it sounds: index columns are the origin of the mysterious unnamed\nfirst column in half the spreadsheets produced by Python.",[119,348,350],{"id":349},"the-transform-expressions-versus-method-chains","The transform: expressions versus method chains",[10,352,353],{},"This is where the two libraries genuinely diverge. pandas mutates and reassigns; Polars describes a\ncomputation and executes it. The Polars version is longer to read the first time and considerably\nharder to get subtly wrong.",[20,355,29,359,29,362,29,365,29,367,29,376,29,382,29,388,29,393,29,397,29,401,29,405,29,408,29,413,29,417,29,420,29,423,29,426,29,429],{"viewBox":22,"role":23,"ariaLabelledBy":356,"xmlns":27,"style":28},[357,358],"pvp-api-t","pvp-api-d",[31,360,361],{"id":357},"Two ways of expressing the same transform",[35,363,364],{"id":358},"pandas builds the result by reassigning frames step by step, while Polars describes one expression graph that it optimises and runs across cores.",[39,366],{"x":41,"y":41,"width":42,"height":43,"fill":44},[39,368],{"x":369,"y":370,"width":371,"height":372,"rx":373,"fill":374,"stroke":375,"style":61},"20.0","26","350.0","164","14","#d9f4f1","var(--teal,#0f9488)",[46,377,381],{"x":378,"y":379,"style":380},"195.0","52","font-size:13px;font-weight:700;fill:var(--teal-ink,#0b6157);text-anchor:middle","pandas",[74,383],{"x1":384,"y1":385,"x2":386,"y2":385,"stroke":375,"style":387},"36.0","62","354.0","stroke-width:1px",[46,389,392],{"x":378,"y":390,"style":391},"84","font-size:11.5px;font-weight:400;fill:var(--text,#172033);text-anchor:middle","assign, then reassign",[46,394,396],{"x":378,"y":395,"style":391},"107","one frame per step",[46,398,400],{"x":378,"y":399,"style":391},"130","single-threaded by default",[46,402,404],{"x":378,"y":403,"style":391},"153","forgiving of messy types",[39,406],{"x":407,"y":370,"width":371,"height":372,"rx":373,"fill":88,"stroke":79,"style":61},"390.0",[46,409,412],{"x":410,"y":379,"style":411},"565.0","font-size:13px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:middle","Polars",[74,414],{"x1":415,"y1":385,"x2":416,"y2":385,"stroke":79,"style":387},"406.0","724.0",[46,418,419],{"x":410,"y":390,"style":391},"one expression graph",[46,421,422],{"x":410,"y":395,"style":391},"no intermediates",[46,424,425],{"x":410,"y":399,"style":391},"parallel across cores",[46,427,428],{"x":410,"y":403,"style":391},"strict about types",[46,430,432],{"x":48,"y":431,"style":116},"218","strictness is the feature, not the friction",[124,434,436],{"className":190,"code":435,"language":192,"meta":129,"style":129},"import polars as pl\n\nsummary = (\n    pl.read_excel(\"sales.xlsx\", sheet_name=\"Raw\")\n      .filter(pl.col(\"Revenue\").is_not_null())\n      .with_columns(\n          (pl.col(\"Revenue\") * 0.2).alias(\"Tax\"),\n          pl.col(\"Region\").str.strip_chars().str.to_uppercase(),\n      )\n      .group_by(\"Region\")\n      .agg(\n          pl.col(\"Revenue\").sum().alias(\"Revenue\"),\n          pl.col(\"Revenue\").mean().round(2).alias(\"Avg deal\"),\n          pl.len().alias(\"Deals\"),\n      )\n      .sort(\"Revenue\", descending=True)\n)\nprint(summary)\n",[131,437,438,448,452,462,479,490,495,520,531,536,546,552,566,586,597,602,622,627],{"__ignoreMap":129},[134,439,440,442,444,446],{"class":74,"line":136},[134,441,200],{"class":199},[134,443,218],{"class":203},[134,445,207],{"class":199},[134,447,223],{"class":203},[134,449,450],{"class":74,"line":213},[134,451,230],{"emptyLinePlaceholder":229},[134,453,454,457,459],{"class":74,"line":226},[134,455,456],{"class":203},"summary ",[134,458,239],{"class":199},[134,460,461],{"class":203}," (\n",[134,463,464,467,469,471,473,475,477],{"class":74,"line":233},[134,465,466],{"class":203},"    pl.read_excel(",[134,468,245],{"class":143},[134,470,248],{"class":203},[134,472,252],{"class":251},[134,474,239],{"class":199},[134,476,257],{"class":143},[134,478,270],{"class":203},[134,480,481,484,487],{"class":74,"line":273},[134,482,483],{"class":203},"      .filter(pl.col(",[134,485,486],{"class":143},"\"Revenue\"",[134,488,489],{"class":203},").is_not_null())\n",[134,491,492],{"class":74,"line":284},[134,493,494],{"class":203},"      .with_columns(\n",[134,496,497,500,502,505,508,511,514,517],{"class":74,"line":289},[134,498,499],{"class":203},"          (pl.col(",[134,501,486],{"class":143},[134,503,504],{"class":203},") ",[134,506,507],{"class":199},"*",[134,509,510],{"class":315}," 0.2",[134,512,513],{"class":203},").alias(",[134,515,516],{"class":143},"\"Tax\"",[134,518,519],{"class":203},"),\n",[134,521,522,525,528],{"class":74,"line":312},[134,523,524],{"class":203},"          pl.col(",[134,526,527],{"class":143},"\"Region\"",[134,529,530],{"class":203},").str.strip_chars().str.to_uppercase(),\n",[134,532,533],{"class":74,"line":322},[134,534,535],{"class":203},"      )\n",[134,537,539,542,544],{"class":74,"line":538},10,[134,540,541],{"class":203},"      .group_by(",[134,543,527],{"class":143},[134,545,270],{"class":203},[134,547,549],{"class":74,"line":548},11,[134,550,551],{"class":203},"      .agg(\n",[134,553,555,557,559,562,564],{"class":74,"line":554},12,[134,556,524],{"class":203},[134,558,486],{"class":143},[134,560,561],{"class":203},").sum().alias(",[134,563,486],{"class":143},[134,565,519],{"class":203},[134,567,569,571,573,576,579,581,584],{"class":74,"line":568},13,[134,570,524],{"class":203},[134,572,486],{"class":143},[134,574,575],{"class":203},").mean().round(",[134,577,578],{"class":315},"2",[134,580,513],{"class":203},[134,582,583],{"class":143},"\"Avg deal\"",[134,585,519],{"class":203},[134,587,589,592,595],{"class":74,"line":588},14,[134,590,591],{"class":203},"          pl.len().alias(",[134,593,594],{"class":143},"\"Deals\"",[134,596,519],{"class":203},[134,598,600],{"class":74,"line":599},15,[134,601,535],{"class":203},[134,603,605,608,610,612,615,617,620],{"class":74,"line":604},16,[134,606,607],{"class":203},"      .sort(",[134,609,486],{"class":143},[134,611,248],{"class":203},[134,613,614],{"class":251},"descending",[134,616,239],{"class":199},[134,618,619],{"class":315},"True",[134,621,270],{"class":203},[134,623,625],{"class":74,"line":624},17,[134,626,270],{"class":203},[134,628,630,632],{"class":74,"line":629},18,[134,631,316],{"class":315},[134,633,634],{"class":203},"(summary)\n",[10,636,637,638,641,642,645,646,649],{},"The same pipeline in pandas needs a ",[131,639,640],{},"dropna",", an assignment per derived column, a ",[131,643,644],{},"groupby().agg()","\nwith a dictionary, and a ",[131,647,648],{},"sort_values"," — and every step produces an intermediate frame. Polars keeps\nit as one expression graph, which is why it parallelises across cores without being asked.",[119,651,653],{"id":652},"writing-a-formatted-workbook-from-polars","Writing a formatted workbook from Polars",[10,655,656,658,659,663],{},[131,657,172],{}," is the part that surprises people coming from pandas: it is a formatting API, not just\na dump. Table styles, per-column number formats, conditional formats and column autofit are all\nkeyword arguments, and underneath it is the same xlsxwriter used by\n",[14,660,662],{"href":661},"\u002Fformatting-and-charting-excel-reports-with-python\u002Fbuilding-excel-reports-with-xlsxwriter\u002F","Building Excel Reports with XlsxWriter",".",[124,665,667],{"className":190,"code":666,"language":192,"meta":129,"style":129},"summary.write_excel(\n    \"regional.xlsx\",\n    worksheet=\"By region\",\n    table_style=\"Table Style Medium 9\",\n    column_formats={\"Revenue\": \"#,##0.00\", \"Avg deal\": \"#,##0.00\"},\n    conditional_formats={\"Revenue\": \"data_bar\"},\n    autofit=True,\n    freeze_panes=\"A2\",\n)\n",[131,668,669,674,682,694,706,735,753,764,776],{"__ignoreMap":129},[134,670,671],{"class":74,"line":136},[134,672,673],{"class":203},"summary.write_excel(\n",[134,675,676,679],{"class":74,"line":213},[134,677,678],{"class":143},"    \"regional.xlsx\"",[134,680,681],{"class":203},",\n",[134,683,684,687,689,692],{"class":74,"line":226},[134,685,686],{"class":251},"    worksheet",[134,688,239],{"class":199},[134,690,691],{"class":143},"\"By region\"",[134,693,681],{"class":203},[134,695,696,699,701,704],{"class":74,"line":233},[134,697,698],{"class":251},"    table_style",[134,700,239],{"class":199},[134,702,703],{"class":143},"\"Table Style Medium 9\"",[134,705,681],{"class":203},[134,707,708,711,713,716,718,721,724,726,728,730,732],{"class":74,"line":273},[134,709,710],{"class":251},"    column_formats",[134,712,239],{"class":199},[134,714,715],{"class":203},"{",[134,717,486],{"class":143},[134,719,720],{"class":203},": ",[134,722,723],{"class":143},"\"#,##0.00\"",[134,725,248],{"class":203},[134,727,583],{"class":143},[134,729,720],{"class":203},[134,731,723],{"class":143},[134,733,734],{"class":203},"},\n",[134,736,737,740,742,744,746,748,751],{"class":74,"line":284},[134,738,739],{"class":251},"    conditional_formats",[134,741,239],{"class":199},[134,743,715],{"class":203},[134,745,486],{"class":143},[134,747,720],{"class":203},[134,749,750],{"class":143},"\"data_bar\"",[134,752,734],{"class":203},[134,754,755,758,760,762],{"class":74,"line":289},[134,756,757],{"class":251},"    autofit",[134,759,239],{"class":199},[134,761,619],{"class":315},[134,763,681],{"class":203},[134,765,766,769,771,774],{"class":74,"line":312},[134,767,768],{"class":251},"    freeze_panes",[134,770,239],{"class":199},[134,772,773],{"class":143},"\"A2\"",[134,775,681],{"class":203},[134,777,778],{"class":74,"line":322},[134,779,270],{"class":203},[10,781,782,783,786,787,790,791,794,795,799],{},"Getting that far in pandas means an ",[131,784,785],{},"ExcelWriter",", a reach through to ",[131,788,789],{},"writer.book",", a format object\nper column and a ",[131,792,793],{},"set_column"," call for each — perhaps fifteen lines to this one. If the deliverable\nis a formatted sheet and the transform is already in Polars, there is no reason to hand back to\npandas to write it. ",[14,796,798],{"href":797},"\u002Fadvanced-data-transformation-and-cleaning\u002Freading-excel-with-polars-and-arrow\u002Fwrite-a-polars-dataframe-to-excel-with-formatting\u002F","Write a Polars DataFrame to Excel with Formatting","\ngoes through the full set of options.",[119,801,803],{"id":802},"moving-between-the-two","Moving between the two",[10,805,806],{},"Neither choice is permanent. Arrow sits underneath both, so a conversion is cheap and often\nzero-copy for numeric columns — which makes \"Polars for the heavy transform, pandas for the one\nlibrary that only speaks DataFrame\" a perfectly reasonable design.",[124,808,810],{"className":190,"code":809,"language":192,"meta":129,"style":129},"import polars as pl\n\npldf = pl.read_excel(\"sales.xlsx\")\npdf = pldf.to_pandas(use_pyarrow_extension_array=True)   # keeps strings on the Arrow side\nback = pl.from_pandas(pdf)\nprint(back.equals(pldf))\n",[131,811,812,822,826,838,861,871],{"__ignoreMap":129},[134,813,814,816,818,820],{"class":74,"line":136},[134,815,200],{"class":199},[134,817,218],{"class":203},[134,819,207],{"class":199},[134,821,223],{"class":203},[134,823,824],{"class":74,"line":213},[134,825,230],{"emptyLinePlaceholder":229},[134,827,828,830,832,834,836],{"class":74,"line":226},[134,829,292],{"class":203},[134,831,239],{"class":199},[134,833,297],{"class":203},[134,835,245],{"class":143},[134,837,270],{"class":203},[134,839,840,842,844,847,850,852,854,857],{"class":74,"line":233},[134,841,236],{"class":203},[134,843,239],{"class":199},[134,845,846],{"class":203}," pldf.to_pandas(",[134,848,849],{"class":251},"use_pyarrow_extension_array",[134,851,239],{"class":199},[134,853,619],{"class":315},[134,855,856],{"class":203},")   ",[134,858,860],{"class":859},"s-wDw","# keeps strings on the Arrow side\n",[134,862,863,866,868],{"class":74,"line":273},[134,864,865],{"class":203},"back ",[134,867,239],{"class":199},[134,869,870],{"class":203}," pl.from_pandas(pdf)\n",[134,872,873,875],{"class":74,"line":284},[134,874,316],{"class":315},[134,876,877],{"class":203},"(back.equals(pldf))\n",[10,879,880,881,883,884,886],{},"The conversion is where the two type systems meet, and it is worth checking rather than assuming: a\npandas ",[131,882,337],{}," column of mixed ints and strings becomes a Polars ",[131,885,333],{}," on the way over, and does\nnot become mixed again on the way back.",[119,888,890],{"id":889},"where-pandas-still-wins","Where pandas still wins",[10,892,893,894,897,898,901,902,905,906,908],{},"Three things keep pandas in Excel work regardless of speed. It has a far larger surface of\nspreadsheet-shaped conveniences — ",[131,895,896],{},"read_excel(sheet_name=None)"," returning a dict of every tab,\n",[131,899,900],{},"pivot_table"," with margins, ",[131,903,904],{},"to_excel"," straight onto an open ",[131,907,785],{},". It is what every other\nlibrary expects: plotting, statistics, machine-learning and reporting packages take DataFrames, and\n\"DataFrame\" means the pandas one in most of them. And its coercion helpers are more forgiving of\ngenuinely dirty data, which is the normal state of a spreadsheet that a person maintains.",[124,910,912],{"className":190,"code":911,"language":192,"meta":129,"style":129},"import pandas as pd\n\n# The messy-column workflow pandas is unusually good at.\nframe = pd.read_excel(\"messy.xlsx\")\nframe[\"Amount\"] = pd.to_numeric(frame[\"Amount\"], errors=\"coerce\")\nframe[\"Date\"] = pd.to_datetime(frame[\"Date\"], errors=\"coerce\", dayfirst=True)\nreport = frame[frame[[\"Amount\", \"Date\"]].isna().any(axis=1)]\nprint(f\"{len(report)} rows need attention\")\n",[131,913,914,924,928,933,947,978,1013,1043],{"__ignoreMap":129},[134,915,916,918,920,922],{"class":74,"line":136},[134,917,200],{"class":199},[134,919,204],{"class":203},[134,921,207],{"class":199},[134,923,210],{"class":203},[134,925,926],{"class":74,"line":213},[134,927,230],{"emptyLinePlaceholder":229},[134,929,930],{"class":74,"line":226},[134,931,932],{"class":859},"# The messy-column workflow pandas is unusually good at.\n",[134,934,935,938,940,942,945],{"class":74,"line":233},[134,936,937],{"class":203},"frame ",[134,939,239],{"class":199},[134,941,242],{"class":203},[134,943,944],{"class":143},"\"messy.xlsx\"",[134,946,270],{"class":203},[134,948,949,952,955,958,960,963,965,968,971,973,976],{"class":74,"line":273},[134,950,951],{"class":203},"frame[",[134,953,954],{"class":143},"\"Amount\"",[134,956,957],{"class":203},"] ",[134,959,239],{"class":199},[134,961,962],{"class":203}," pd.to_numeric(frame[",[134,964,954],{"class":143},[134,966,967],{"class":203},"], ",[134,969,970],{"class":251},"errors",[134,972,239],{"class":199},[134,974,975],{"class":143},"\"coerce\"",[134,977,270],{"class":203},[134,979,980,982,985,987,989,992,994,996,998,1000,1002,1004,1007,1009,1011],{"class":74,"line":284},[134,981,951],{"class":203},[134,983,984],{"class":143},"\"Date\"",[134,986,957],{"class":203},[134,988,239],{"class":199},[134,990,991],{"class":203}," pd.to_datetime(frame[",[134,993,984],{"class":143},[134,995,967],{"class":203},[134,997,970],{"class":251},[134,999,239],{"class":199},[134,1001,975],{"class":143},[134,1003,248],{"class":203},[134,1005,1006],{"class":251},"dayfirst",[134,1008,239],{"class":199},[134,1010,619],{"class":315},[134,1012,270],{"class":203},[134,1014,1015,1018,1020,1023,1025,1027,1029,1032,1035,1037,1040],{"class":74,"line":289},[134,1016,1017],{"class":203},"report ",[134,1019,239],{"class":199},[134,1021,1022],{"class":203}," frame[frame[[",[134,1024,954],{"class":143},[134,1026,248],{"class":203},[134,1028,984],{"class":143},[134,1030,1031],{"class":203},"]].isna().any(",[134,1033,1034],{"class":251},"axis",[134,1036,239],{"class":199},[134,1038,1039],{"class":315},"1",[134,1041,1042],{"class":203},")]\n",[134,1044,1045,1047,1050,1053,1056,1059,1062,1065,1068,1071],{"class":74,"line":312},[134,1046,316],{"class":315},[134,1048,1049],{"class":203},"(",[134,1051,1052],{"class":199},"f",[134,1054,1055],{"class":143},"\"",[134,1057,715],{"class":1058},"sSjpA",[134,1060,1061],{"class":315},"len",[134,1063,1064],{"class":203},"(report)",[134,1066,1067],{"class":1058},"}",[134,1069,1070],{"class":143}," rows need attention\"",[134,1072,270],{"class":203},[10,1074,1075,1076,1079,1080,1083,1084,663],{},"Polars can do all of that, with ",[131,1077,1078],{},"strict=False"," casts and ",[131,1081,1082],{},"str.to_date",", but it will make you say\nwhat you mean at each step. That is the right trade once the data is clean and an obstacle while\nyou are still finding out what is in it — the workflow described in\n",[14,1085,1087],{"href":1086},"\u002Fadvanced-data-transformation-and-cleaning\u002Fcleaning-excel-data-with-pandas\u002F","Cleaning Excel Data with Pandas",[119,1089,1091],{"id":1090},"multi-sheet-workbooks-in-both-libraries","Multi-sheet workbooks in both libraries",[10,1093,1094],{},"Reporting workbooks rarely have one tab, and the two libraries handle that differently enough to\nmatter. pandas returns a dictionary keyed by sheet name; Polars returns one too, but its reader\ntakes a list of sheet names or indices directly and its frames carry no index to reconcile\nafterwards.",[124,1096,1098],{"className":190,"code":1097,"language":192,"meta":129,"style":129},"import pandas as pd\nimport polars as pl\n\npandas_tabs = pd.read_excel(\"book.xlsx\", sheet_name=None)          # {name: DataFrame}\ncombined_pd = pd.concat(\n    [frame.assign(Source=name) for name, frame in pandas_tabs.items()],\n    ignore_index=True,\n)\n\npolars_tabs = pl.read_excel(\"book.xlsx\", sheet_name=[\"Jan\", \"Feb\", \"Mar\"])\ncombined_pl = pl.concat(\n    [frame.with_columns(pl.lit(name).alias(\"Source\")) for name, frame in polars_tabs.items()],\n    how=\"diagonal\",\n)\nprint(combined_pd.shape, combined_pl.shape)\n",[131,1099,1100,1110,1120,1124,1151,1161,1186,1197,1201,1205,1241,1251,1271,1283,1287],{"__ignoreMap":129},[134,1101,1102,1104,1106,1108],{"class":74,"line":136},[134,1103,200],{"class":199},[134,1105,204],{"class":203},[134,1107,207],{"class":199},[134,1109,210],{"class":203},[134,1111,1112,1114,1116,1118],{"class":74,"line":213},[134,1113,200],{"class":199},[134,1115,218],{"class":203},[134,1117,207],{"class":199},[134,1119,223],{"class":203},[134,1121,1122],{"class":74,"line":226},[134,1123,230],{"emptyLinePlaceholder":229},[134,1125,1126,1129,1131,1133,1136,1138,1140,1142,1145,1148],{"class":74,"line":233},[134,1127,1128],{"class":203},"pandas_tabs ",[134,1130,239],{"class":199},[134,1132,242],{"class":203},[134,1134,1135],{"class":143},"\"book.xlsx\"",[134,1137,248],{"class":203},[134,1139,252],{"class":251},[134,1141,239],{"class":199},[134,1143,1144],{"class":315},"None",[134,1146,1147],{"class":203},")          ",[134,1149,1150],{"class":859},"# {name: DataFrame}\n",[134,1152,1153,1156,1158],{"class":74,"line":273},[134,1154,1155],{"class":203},"combined_pd ",[134,1157,239],{"class":199},[134,1159,1160],{"class":203}," pd.concat(\n",[134,1162,1163,1166,1169,1171,1174,1177,1180,1183],{"class":74,"line":284},[134,1164,1165],{"class":203},"    [frame.assign(",[134,1167,1168],{"class":251},"Source",[134,1170,239],{"class":199},[134,1172,1173],{"class":203},"name) ",[134,1175,1176],{"class":199},"for",[134,1178,1179],{"class":203}," name, frame ",[134,1181,1182],{"class":199},"in",[134,1184,1185],{"class":203}," pandas_tabs.items()],\n",[134,1187,1188,1191,1193,1195],{"class":74,"line":289},[134,1189,1190],{"class":251},"    ignore_index",[134,1192,239],{"class":199},[134,1194,619],{"class":315},[134,1196,681],{"class":203},[134,1198,1199],{"class":74,"line":312},[134,1200,270],{"class":203},[134,1202,1203],{"class":74,"line":322},[134,1204,230],{"emptyLinePlaceholder":229},[134,1206,1207,1210,1212,1214,1216,1218,1220,1222,1225,1228,1230,1233,1235,1238],{"class":74,"line":538},[134,1208,1209],{"class":203},"polars_tabs ",[134,1211,239],{"class":199},[134,1213,297],{"class":203},[134,1215,1135],{"class":143},[134,1217,248],{"class":203},[134,1219,252],{"class":251},[134,1221,239],{"class":199},[134,1223,1224],{"class":203},"[",[134,1226,1227],{"class":143},"\"Jan\"",[134,1229,248],{"class":203},[134,1231,1232],{"class":143},"\"Feb\"",[134,1234,248],{"class":203},[134,1236,1237],{"class":143},"\"Mar\"",[134,1239,1240],{"class":203},"])\n",[134,1242,1243,1246,1248],{"class":74,"line":548},[134,1244,1245],{"class":203},"combined_pl ",[134,1247,239],{"class":199},[134,1249,1250],{"class":203}," pl.concat(\n",[134,1252,1253,1256,1259,1262,1264,1266,1268],{"class":74,"line":554},[134,1254,1255],{"class":203},"    [frame.with_columns(pl.lit(name).alias(",[134,1257,1258],{"class":143},"\"Source\"",[134,1260,1261],{"class":203},")) ",[134,1263,1176],{"class":199},[134,1265,1179],{"class":203},[134,1267,1182],{"class":199},[134,1269,1270],{"class":203}," polars_tabs.items()],\n",[134,1272,1273,1276,1278,1281],{"class":74,"line":568},[134,1274,1275],{"class":251},"    how",[134,1277,239],{"class":199},[134,1279,1280],{"class":143},"\"diagonal\"",[134,1282,681],{"class":203},[134,1284,1285],{"class":74,"line":588},[134,1286,270],{"class":203},[134,1288,1289,1291],{"class":74,"line":599},[134,1290,316],{"class":315},[134,1292,1293],{"class":203},"(combined_pd.shape, combined_pl.shape)\n",[10,1295,1296,1299,1300,1303,1304,1307,1308,1312],{},[131,1297,1298],{},"how=\"diagonal\""," is the detail worth keeping: it unions tabs whose columns do not match exactly,\nfilling the gaps with nulls, where a plain ",[131,1301,1302],{},"pl.concat"," would raise. pandas' ",[131,1305,1306],{},"concat"," does the same\nthing silently, which is friendlier until a renamed column in one month's tab produces two columns\nwhere you expected one. The trade is the same one that runs through this whole comparison — the\nstricter library tells you about the problem, and the forgiving one hands you a result that looks\nfine. ",[14,1309,1311],{"href":1310},"\u002Fadvanced-data-transformation-and-cleaning\u002Fmerging-and-joining-excel-dataframes\u002Fconcatenate-excel-sheets-with-different-columns\u002F","Concatenate Excel Sheets with Different Columns","\ncovers the failure modes of that union in detail.",[119,1314,1316],{"id":1315},"common-pitfalls","Common pitfalls",[1318,1319,1320,1336],"table",{},[1321,1322,1323],"thead",{},[1324,1325,1326,1330,1333],"tr",{},[1327,1328,1329],"th",{},"Symptom",[1327,1331,1332],{},"Cause",[1327,1334,1335],{},"Fix",[1337,1338,1339,1362,1386,1399,1415],"tbody",{},[1324,1340,1341,1347,1350],{},[1342,1343,1344],"td",{},[131,1345,1346],{},"ModuleNotFoundError: fastexcel",[1342,1348,1349],{},"Polars' Excel reader is an optional extra",[1342,1351,1352,1355,1356,1359,1360],{},[131,1353,1354],{},"pip install fastexcel",", or pass ",[131,1357,1358],{},"engine=\"openpyxl\""," to ",[131,1361,176],{},[1324,1363,1364,1369,1375],{},[1342,1365,1366,1367],{},"Polars read returns everything as ",[131,1368,333],{},[1342,1370,1371,1374],{},[131,1372,1373],{},"infer_schema_length"," too small, or a header row of mixed types",[1342,1376,1377,1378,1381,1382,1385],{},"Pass ",[131,1379,1380],{},"infer_schema_length=None"," to scan the whole column, or a ",[131,1383,1384],{},"schema_overrides"," dict",[1324,1387,1388,1393,1396],{},[1342,1389,1390,1392],{},[131,1391,172],{}," raises on an existing file",[1342,1394,1395],{},"It builds a new workbook; it cannot open one",[1342,1397,1398],{},"Write a new file, or use openpyxl if the target must be edited in place",[1324,1400,1401,1404,1407],{},[1342,1402,1403],{},"A conversion to pandas is slower than the transform",[1342,1405,1406],{},"Object-dtype strings are copied element by element",[1342,1408,1377,1409,1359,1412],{},[131,1410,1411],{},"use_pyarrow_extension_array=True",[131,1413,1414],{},"to_pandas()",[1324,1416,1417,1420,1423],{},[1342,1418,1419],{},"Dates arrive as integers",[1342,1421,1422],{},"Excel serial numbers were read without date inference",[1342,1424,1425,1426],{},"Cast explicitly, as in ",[14,1427,1429],{"href":1428},"\u002Fadvanced-data-transformation-and-cleaning\u002Fworking-with-dates-and-times-in-excel-data\u002Ffix-excel-serial-numbers-showing-instead-of-dates\u002F","Fix Excel Serial Numbers Showing Instead of Dates",[119,1431,1433],{"id":1432},"performance-and-scale","Performance and scale",[20,1435,29,1441,29,1444,29,1447,29,1451,29,1457,29,1463,29,1470,29,1476,29,1480,29,1482,29,1486,29,1491,29,1495,29,1498,29,1504,29,1509,29,1513,29,1516,29,1520,29,1525,29,1529],{"viewBox":1436,"role":23,"ariaLabelledBy":1437,"xmlns":27,"style":1440},"0 0 720 240",[1438,1439],"pvp-scale-t","pvp-scale-d","width:100%;max-width:720px;height:auto;display:block;margin:1.5rem auto;font-family:Inter,ui-sans-serif,system-ui,sans-serif",[31,1442,1443],{"id":1438},"Relative time on a one-million-row group-and-sort",[35,1445,1446],{"id":1439},"Reading is nearly identical when both libraries use calamine. The grouping and sorting step is several times faster in Polars, and a lazy scan over a Parquet copy is faster again.",[39,1448],{"x":41,"y":41,"width":1449,"height":1450,"fill":44},"720","240",[46,1452,1456],{"x":1453,"y":1454,"style":1455},"20","56","font-size:12px;font-weight:600;fill:var(--text,#172033);text-anchor:start","pandas transform",[39,1458],{"x":1459,"y":1460,"width":1461,"height":370,"rx":1462,"fill":59,"stroke":60},"200","40","344.0","6",[39,1464],{"x":1465,"y":1466,"width":1467,"height":1468,"rx":1469,"fill":374,"stroke":375},"201","41","342.0","24","5",[46,1471,1475],{"x":1472,"y":1473,"style":1474},"556.0","58","font-size:12px;font-weight:700;fill:var(--teal-ink,#0b6157);text-anchor:start","baseline",[46,1477,1479],{"x":1453,"y":1478,"style":1455},"100","Polars transform",[39,1481],{"x":1459,"y":390,"width":1461,"height":370,"rx":1462,"fill":59,"stroke":60},[39,1483],{"x":1465,"y":1484,"width":1485,"height":1468,"rx":1469,"fill":88,"stroke":79},"85","79.3",[46,1487,1490],{"x":1472,"y":1488,"style":1489},"102","font-size:12px;font-weight:700;fill:var(--brand-strong,#4338ca);text-anchor:start","multi-threaded",[46,1492,1494],{"x":1453,"y":1493,"style":1455},"144","Polars lazy scan",[39,1496],{"x":1459,"y":1497,"width":1461,"height":370,"rx":1462,"fill":59,"stroke":60},"128",[39,1499],{"x":1465,"y":1500,"width":1501,"height":1468,"rx":1469,"fill":1502,"stroke":1503},"129","28","#fdefd8","var(--gold,#b4740a)",[46,1505,1508],{"x":1472,"y":1506,"style":1507},"146","font-size:12px;font-weight:700;fill:var(--gold-ink,#7a4e06);text-anchor:start","columns pruned",[46,1510,1512],{"x":1453,"y":1511,"style":1455},"188","shared calamine read",[39,1514],{"x":1459,"y":1515,"width":1461,"height":370,"rx":1462,"fill":59,"stroke":60},"172",[39,1517],{"x":1465,"y":1518,"width":1519,"height":1468,"rx":1469,"fill":59,"stroke":60},"173","186.4",[46,1521,1524],{"x":1472,"y":1522,"style":1523},"190","font-size:12px;font-weight:700;fill:var(--muted,#5b6780);text-anchor:start","identical either way",[46,1526,1528],{"x":1453,"y":1453,"style":1527},"font-size:11.5px;font-weight:600;fill:var(--muted,#5b6780);text-anchor:start","relative cost",[46,1530,1533],{"x":1531,"y":1532,"style":116},"360.0","230","the read is shared ground; the transform is not",[10,1535,1536],{},"The read is a wash when both libraries use calamine, because they are running the same Rust parser.\nThe transform is where Polars pulls ahead, and the gap widens with row count: it is multi-threaded\nby default and never materialises the intermediate frames a pandas chain produces. Memory follows\nthe same shape — Arrow buffers are more compact than NumPy object arrays for strings, which is the\ncolumn type Excel exports produce most.",[10,1538,1539,1540,1543],{},"For very large inputs Polars has one more gear that pandas has no equivalent of. ",[131,1541,1542],{},"scan_parquet","\nbuilds a lazy query and only reads the columns and rows the query needs, so a pipeline that starts\nby converting the workbook once can then run against a fraction of the data:",[124,1545,1547],{"className":190,"code":1546,"language":192,"meta":129,"style":129},"import polars as pl\n\npl.read_excel(\"huge.xlsx\").write_parquet(\"huge.parquet\")     # once\n\nresult = (\n    pl.scan_parquet(\"huge.parquet\")\n      .filter(pl.col(\"Region\") == \"North\")\n      .group_by(\"Rep\").agg(pl.col(\"Revenue\").sum())\n      .collect()\n)\n",[131,1548,1549,1559,1563,1583,1587,1596,1605,1621,1636,1641],{"__ignoreMap":129},[134,1550,1551,1553,1555,1557],{"class":74,"line":136},[134,1552,200],{"class":199},[134,1554,218],{"class":203},[134,1556,207],{"class":199},[134,1558,223],{"class":203},[134,1560,1561],{"class":74,"line":213},[134,1562,230],{"emptyLinePlaceholder":229},[134,1564,1565,1568,1571,1574,1577,1580],{"class":74,"line":226},[134,1566,1567],{"class":203},"pl.read_excel(",[134,1569,1570],{"class":143},"\"huge.xlsx\"",[134,1572,1573],{"class":203},").write_parquet(",[134,1575,1576],{"class":143},"\"huge.parquet\"",[134,1578,1579],{"class":203},")     ",[134,1581,1582],{"class":859},"# once\n",[134,1584,1585],{"class":74,"line":233},[134,1586,230],{"emptyLinePlaceholder":229},[134,1588,1589,1592,1594],{"class":74,"line":273},[134,1590,1591],{"class":203},"result ",[134,1593,239],{"class":199},[134,1595,461],{"class":203},[134,1597,1598,1601,1603],{"class":74,"line":284},[134,1599,1600],{"class":203},"    pl.scan_parquet(",[134,1602,1576],{"class":143},[134,1604,270],{"class":203},[134,1606,1607,1609,1611,1613,1616,1619],{"class":74,"line":289},[134,1608,483],{"class":203},[134,1610,527],{"class":143},[134,1612,504],{"class":203},[134,1614,1615],{"class":199},"==",[134,1617,1618],{"class":143}," \"North\"",[134,1620,270],{"class":203},[134,1622,1623,1625,1628,1631,1633],{"class":74,"line":312},[134,1624,541],{"class":203},[134,1626,1627],{"class":143},"\"Rep\"",[134,1629,1630],{"class":203},").agg(pl.col(",[134,1632,486],{"class":143},[134,1634,1635],{"class":203},").sum())\n",[134,1637,1638],{"class":74,"line":322},[134,1639,1640],{"class":203},"      .collect()\n",[134,1642,1643],{"class":74,"line":538},[134,1644,270],{"class":203},[10,1646,1647,1648,663],{},"Excel itself has no lazy mode, so this only pays off when the same workbook is queried repeatedly —\nthe conversion argument made in ",[14,1649,1651],{"href":1650},"\u002Fadvanced-data-transformation-and-cleaning\u002Freading-excel-with-polars-and-arrow\u002Fconvert-excel-files-to-parquet-with-python\u002F","Convert Excel Files to Parquet with Python",[119,1653,1655],{"id":1654},"conclusion","Conclusion",[10,1657,1658],{},"Choose Polars for the middle of the pipeline and pandas for the edges of the ecosystem. Polars is\nfaster and stricter on transforms, has a genuinely better formatted-Excel writer through xlsxwriter,\nand scales further before memory becomes the constraint. pandas remains the language every other\nlibrary speaks and the more forgiving tool while data is still messy. Because Arrow sits under both,\nthe decision is reversible in a single line, so it is worth making per pipeline rather than per team.",[119,1660,1662],{"id":1661},"frequently-asked-questions","Frequently asked questions",[10,1664,1665,1669],{},[1666,1667,1668],"strong",{},"Does Polars read Excel without pandas installed?","\nYes. polars.read_excel() calls python-calamine (or openpyxl, or xlsx2csv) directly and returns a Polars DataFrame. pandas is not involved and does not need to be installed.",[10,1671,1672,1675],{},[1666,1673,1674],{},"Can I convert between the two without copying the data?","\nMostly. df.to_pandas() and pl.from_pandas() go through Arrow, which is zero-copy for numeric and boolean columns and a real copy for Python object strings. Passing use_pyarrow_extension_array=True keeps strings on the Arrow side too.",[10,1677,1678,1681],{},[1666,1679,1680],{},"Does Polars write formatted Excel?","\nYes — write_excel wraps xlsxwriter, so it accepts table styles, column formats, conditional formats and autofit in one call. It cannot edit an existing workbook, because xlsxwriter cannot.",[10,1683,1684,1687],{},[1666,1685,1686],{},"Is Polars always faster?","\nFor the transform, usually, and by more as the data grows. For the Excel read itself the parser decides, and both libraries can use calamine — so a like-for-like read is close. On a 2,000-row monthly report neither is measurably faster.",[10,1689,1690,1693],{},[1666,1691,1692],{},"Which one handles messy spreadsheet data better?","\npandas, on balance. Its coercion helpers — to_numeric with errors='coerce', fillna, interpolate — are more forgiving of the sort of mixed-type column an Excel export produces. Polars is stricter, which is a virtue once the data is clean and a friction before that.",[119,1695,1697],{"id":1696},"related","Related",[1699,1700,1701,1708,1715,1722,1729],"ul",{},[1702,1703,1704,1705,1707],"li",{},"Up one level: ",[14,1706,17],{"href":16}," — the full landscape including openpyxl, xlsxwriter and xlwings.",[1702,1709,1710,1714],{},[14,1711,1713],{"href":1712},"\u002Fadvanced-data-transformation-and-cleaning\u002Freading-excel-with-polars-and-arrow\u002F","Reading Excel with Polars and Arrow"," — the Polars-side topic, from reads to Parquet conversion.",[1702,1716,1717,1721],{},[14,1718,1720],{"href":1719},"\u002Fadvanced-data-transformation-and-cleaning\u002Freading-excel-with-polars-and-arrow\u002Fread-an-excel-file-with-polars-read-excel\u002F","Read an Excel File with polars.read_excel"," — engine choice, schemas and multi-sheet reads in detail.",[1702,1723,1724,1728],{},[14,1725,1727],{"href":1726},"\u002Fgetting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002Fopenpyxl-vs-pandas-for-excel-automation\u002F","openpyxl vs pandas for Excel Automation"," — the other half of the library question.",[1702,1730,1731,1735],{},[14,1732,1734],{"href":1733},"\u002Fadvanced-data-transformation-and-cleaning\u002Freading-excel-with-polars-and-arrow\u002Fspeed-up-pandas-excel-reads-with-the-calamine-engine\u002F","Speed Up pandas Excel Reads with the calamine Engine"," — the parser both libraries can share.",[1737,1738,1739],"style",{},"html pre.shiki code .sMTad, html code.shiki .sMTad{--shiki-default:#6F42C1;--shiki-dark:#FFB757}html pre.shiki code .srMev, html code.shiki .srMev{--shiki-default:#032F62;--shiki-dark:#ADDCFF}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .s-kum, html code.shiki .s-kum{--shiki-default:#D73A49;--shiki-dark:#FF9492}html pre.shiki code .skGVy, html code.shiki .skGVy{--shiki-default:#24292E;--shiki-dark:#F0F3F6}html pre.shiki code .sa561, html code.shiki .sa561{--shiki-default:#E36209;--shiki-dark:#FFB757}html pre.shiki code .sP0c6, html code.shiki .sP0c6{--shiki-default:#005CC5;--shiki-dark:#91CBFF}html pre.shiki code .s-wDw, html code.shiki .s-wDw{--shiki-default:#6A737D;--shiki-dark:#BDC4CC}html pre.shiki code .sSjpA, html code.shiki .sSjpA{--shiki-default:#005CC5;--shiki-dark:#FF9492}",{"title":129,"searchDepth":213,"depth":213,"links":1741},[1742,1743,1744,1745,1746,1747,1748,1749,1750,1751,1752,1753],{"id":121,"depth":213,"text":122},{"id":183,"depth":213,"text":184},{"id":349,"depth":213,"text":350},{"id":652,"depth":213,"text":653},{"id":802,"depth":213,"text":803},{"id":889,"depth":213,"text":890},{"id":1090,"depth":213,"text":1091},{"id":1315,"depth":213,"text":1316},{"id":1432,"depth":213,"text":1433},{"id":1654,"depth":213,"text":1655},{"id":1661,"depth":213,"text":1662},{"id":1696,"depth":213,"text":1697},"2026-09-04","Both read Excel through calamine and write through xlsxwriter, so the real difference is the transform. Compare the APIs, the speed, and how to move between them.","md",[1758,1760,1762,1764,1766],{"q":1668,"a":1759},"Yes. polars.read_excel() calls python-calamine (or openpyxl, or xlsx2csv) directly and returns a Polars DataFrame. pandas is not involved and does not need to be installed.",{"q":1674,"a":1761},"Mostly. df.to_pandas() and pl.from_pandas() go through Arrow, which is zero-copy for numeric and boolean columns and a real copy for Python object strings. Passing use_pyarrow_extension_array=True keeps strings on the Arrow side too.",{"q":1680,"a":1763},"Yes — write_excel wraps xlsxwriter, so it accepts table styles, column formats, conditional formats and autofit in one call. It cannot edit an existing workbook, because xlsxwriter cannot.",{"q":1686,"a":1765},"For the transform, usually, and by more as the data grows. For the Excel read itself the parser decides, and both libraries can use calamine — so a like-for-like read is close. On a 2,000-row monthly report neither is measurably faster.",{"q":1692,"a":1767},"pandas, on balance. Its coercion helpers — to_numeric with errors='coerce', fillna, interpolate — are more forgiving of the sort of mixed-type column an Excel export produces. Polars is stricter, which is a virtue once the data is clean and a friction before that.",{"breadcrumb":1769},[1770,1772,1775],{"name":1771,"item":177},"Home",{"name":1773,"item":1774},"Getting Started with Python Excel Automation","\u002Fgetting-started-with-python-excel-automation\u002F",{"name":17,"item":16},"\u002Fgetting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002Fpandas-vs-polars-for-excel-workflows",{"title":5,"description":1778},"Compare pandas and Polars for Excel work: shared parsers, expression API versus method chains, formatted write_excel output, conversions, and where each still wins.","pandas-vs-polars-for-excel-workflows","getting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002Fpandas-vs-polars-for-excel-workflows\u002Findex","how-to","HDmMpBSffWO_4PnLmVVwbj2NF3j44mopfgWTsh9G4Z4",[1784,1787],{"title":1727,"path":1785,"stem":1786,"children":-1},"\u002Fgetting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002Fopenpyxl-vs-pandas-for-excel-automation","getting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002Fopenpyxl-vs-pandas-for-excel-automation\u002Findex",{"title":1788,"path":1789,"stem":1790,"children":-1},"Pick an Excel Engine for .xlsx, .xlsm, .xls, .xlsb and .ods","\u002Fgetting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002Fpick-an-excel-engine-for-xlsx-xlsm-xls-xlsb-and-ods","getting-started-with-python-excel-automation\u002Fchoosing-a-python-excel-library\u002Fpick-an-excel-engine-for-xlsx-xlsm-xls-xlsb-and-ods\u002Findex",1788710154656]