

# Release Notes
<a name="data-wrangler-release-notes"></a>

Data Wrangler is regularly updated with new features and bug fixes. To upgrade the version of Data Wrangler you are using in Studio Classic, follow the instructions in [Shut Down and Update Amazon SageMaker Studio Classic Apps](studio-tasks-update-apps.md).



| Release Notes | 
| --- | 
| **8/31/2023**<br />New functionality:<br />You can now create a Data Quality and Insights report on your entire dataset. For more information, see [Get Insights On Data and Data Quality](data-wrangler-data-insights.md).<br />**5/20/2023**<br />New functionality:<br />You can now import your data from Salesforce Data Cloud. For more information, see [Import data from Salesforce Data Cloud](data-wrangler-import.md#data-wrangler-import-salesforce-data-cloud).<br />**4/18/2023**<br />New functionality:<br />You can now get your data in a format that Amazon Personalize can interpret. For more information, see [Map Columns for Amazon Personalize](data-wrangler-transform.md#data-wrangler-transform-personalize).<br />**3/1/2023**<br />New functionality:<br />You can now use Hive to import your data from Amazon EMR. For more information, see [Import data from Amazon EMR](data-wrangler-import.md#data-wrangler-emr).<br />**12/10/2022**<br />New functionality:<br />You can now export your Data Wrangler flow to an inference endpoint. For more information, see [Export to an Inference Endpoint](data-wrangler-data-export.md#data-wrangler-data-export-inference).<br />New functionality:<br />You can now use an interactive notebook widget for data preparation. For more information, see [Use an Interactive Data Preparation Widget in an Amazon SageMaker Studio Classic Notebook to Get Data Insights](data-wrangler-interactively-prepare-data-notebook.md).<br />New functionality:<br />You can now import data from SaaS platforms. For more information, see [Import Data From Software as a Service (SaaS) Platforms](data-wrangler-import.md#data-wrangler-import-saas).<br />**10/12/2022**<br />New functionality:<br />You can now reuse data flows for different data sets. For more information, see [Reusing Data Flows for Different Datasets](data-wrangler-parameterize.md).<br />**10/05/2022**<br />New functionality:<br />You can now use Principal Component Analysis (PCA) as a transform. For more information, see [Reduce Dimensionality within a Dataset](data-wrangler-transform.md#data-wrangler-transform-dimensionality-reduction).<br />**10/05/2022**<br />New functionality:<br />You can now refit parameters in your Data Wrangler flow. For more information, see [Export](data-wrangler-data-export.md).<br />**10/03/2022**<br />New functionality:<br />You can now deploy models from your Data Wrangler flow. For more information, see [Automatically Train Models on Your Data Flow](data-wrangler-autopilot.md).<br />**9/20/2022**<br />New functionality:<br />You can now set data retention periods in Athena. For more information, see [Import data from Athena](data-wrangler-import.md#data-wrangler-import-athena).<br />**6/9/2022**<br />New functionality:<br />You can now use Amazon SageMaker Autopilot to train a model directly from your Data Wrangler flow. For more information, see [Automatically Train Models on Your Data Flow](data-wrangler-autopilot.md).<br />**5/6/2022**<br />New functionality:<br />You can now use additional m5 and r5 instances. For more information, see [Instances](data-wrangler-data-flow.md#data-wrangler-data-flow-instances).<br />**4/27/2022**<br />New functionalities:+  You can now get a data quality report. For more information, see [Get Insights On Data and Data Quality](data-wrangler-data-insights.md) <br />+  You can now perform random sampling and stratified sampling. For more information, see [Sampling](data-wrangler-transform.md#data-wrangler-transform-sampling). <br />**4/1/2022**<br />New functionality:<br />You can now use Databricks as a data source. For more information, see [Import data from Databricks (JDBC)](data-wrangler-import.md#data-wrangler-databricks).<br />**2/2/2022**<br />New functionalities:+  You can now export using destination nodes. For more information, see [Export](data-wrangler-data-export.md) <br />+  You can import ORC and JSON files. For more information about file types, see [Import](data-wrangler-import.md). <br />+  Data Wrangler now supports using the SMOTE transform. For more information, see [Balance Data](data-wrangler-transform.md#data-wrangler-transform-balance-data). <br />+  Data Wrangler now supports similarity encoding for categorical data. For more information, see [Similarity encode](data-wrangler-transform.md#data-wrangler-transform-cat-encode-similarity). <br />+  Data Wrangler now supports unnesting JSON data. For more information, see [Unnest JSON Data](data-wrangler-transform.md#data-wrangler-transform-flatten-column). <br />+  Data Wrangler now supports expanding the values of an array into separate columns. For more information, see [Explode Array](data-wrangler-transform.md#data-wrangler-transform-explode-array). <br />+  Data Wrangler now supports reaching out to the service team when you're having issues. For more information, see [Troubleshoot](data-wrangler-trouble-shooting.md). <br />+  Data Wrangler supports editing and deleting steps in your data flow. For more information, see [Delete a Step from Your Data Flow](data-wrangler-data-flow.md#data-wrangler-data-flow-delete-step) and [Edit a Step in Your Data Wrangler Flow](data-wrangler-data-flow.md#data-wrangler-data-flow-edit-step). <br />+  You can now perform transformations on multiple columns. For more information, see [Transform Data](data-wrangler-transform.md). <br />+  Data Wrangler now supports cost allocation tags. For more information, see [Using Cost Allocation Tags](https://docs.aws.amazon.com/awsaccountbilling/latest/aboutv2/cost-alloc-tags.html). <br />**10/16/2021**<br />New functionality:<br />Data Wrangler now supports Athena workgroups. For more information, see [Import data from Athena](data-wrangler-import.md#data-wrangler-import-athena).<br />**10/6/2021**<br />New functionality:<br />Data Wrangler now supports transforming time series data. For more information, see [Transform Time Series](data-wrangler-transform.md#data-wrangler-transform-time-series).<br />**7/15/2021**<br />New functionalities:+  [Snowflake and Data Wrangler](data-wrangler-security.md#data-wrangler-security-snowflake) is now supported. You can use Snowflake as a data source in Data Wrangler. <br />+  Added support for custom field delimiter in CSV. Now comma, colon, semicolon, pipe (\|) and Tab are supported. <br />+  Now you can export results directly to Amazon S3. <br />+  Added a few new multicollinearity analyzers: Variance Inflation Factors, Principal Component Analysis and Lasso feature selection.  <br />Enhancements:+  The analyze charts can no longer be could be packed with overlapping labels. <br /> Bug Fixes:+  One-hot encoder handles empty string gracefully. <br />+  Fixed crashes that occured when a dataframe column name contained dots.  <br />**4/26/2021** <br />Enhancements:+  Added support for distributed processing Jobs. You can use multiple instances when running a processing job. <br />+  Data Wrangler Processing job now automatically coalesces small outputs when estimated result size is less than 1 gigabytes. <br />+  Feature Store Notebook: Improved feature store ingestion performance <br />+  Data Wrangler Processing jobs now use 1.x as the authoritative container tag for future releases. <br />Bug Fixes:+  Fixed rendering issues for faceted histogram. <br />+  Fixed **Export to Processing Job** to support vector type columns. <br />+  Fixed `Extract using regex` operator to return the first captured group if one or more exists in the regular expression or regex. <br />**2/8/2021** <br />New Functionalities:+  Data Wrangler Flows supports multiple instances. <br />+  Updated Export to Data Wrangler Job Notebook to use SageMaker SDK 2.20.0. <br />+  Updated Export to Pipeline Notebook to use SageMaker SDK 2.20.0. <br />+  Updated Export to Pipeline Notebook to add XGBoost training example as an optional step. <br />Enhancements:+  To improve performance, importing CSV files that contain multiple lines in a single field is no longer supported. <br />Bug Fixes:+  Fixed type inference issue in Quick model.  <br />+  Fixed the bias metric bug in bias reports. <br />+  Fixed the Featurize text transform to work with columns with missing values. <br />+  Fixed Histogram and Scatter plot built-in visualizations to work with datasets that contain array-like columns. <br />+  Athena query now re-runs if the query execution ID has expired.  | 