Clean Dataset
AI-generatedSummary
Clean and normalize a tabular JSON dataset by optionally normalizing column names, trimming and formatting string values, handling configured null-like values, removing currency symbols, converting European formatted numbers, removing empty columns, and removing duplicate rows according to specified deduplication rules.
Inputs
- Normalize Column Names — Boolean to enable normalization of column names before cleaning values (default true).
- Column Name Style — Formatting style for normalized column names: camelCase, lower case with spaces, or snake_case (default snake_case).
- Trim Strings — Boolean to trim leading and trailing whitespace from string values (default true).
- Collapse Repeated Whitespace — Boolean to collapse repeated internal whitespace in strings to a single space (default true).
- String Case — Optionally convert string values to lowercase, uppercase, or preserve original case (default preserve).
- Treat Configured Nulls as Null — Boolean to treat certain user-configured string tokens as null-like values (default true).
- Null Values — Comma-separated list of string tokens to treat as null-like values (default ',NULL,null,N/A,NA,-').
- Null Replacement — Value to use when replacing null-like values: empty string, null, or zero (default null).
- Clean Currency Symbols — Boolean to remove currency symbols from number-like string values (default false).
- Convert European Numbers — Boolean to convert European formatted numbers such as '1.234,56' to standard numeric format (default false).
- Numeric Columns — Comma-separated list of columns for which numeric conversion should be forced when converting European numbers (optional).
- Remove Empty Columns — Boolean to remove columns that contain only null-like values (default false).
- Remove Duplicates — Boolean to remove duplicate rows from the dataset (default false).
- Deduplicate By — Method of identifying duplicates: by full row or by configured keys (default full row).
- Deduplication Keys — Comma-separated list of column names to identify duplicates when 'Deduplicate By' is set to keys (optional).
Output shape
a single aggregated JSON object representing the cleaned dataset and cleaning metadata
The output JSON includes cleaned rows, a map of original to normalized column names, details of removed columns and duplicates, a summary of cleaning actions performed, warnings and errors encountered, and an audit trail of cleaning steps.
Examples
Example 1: Basic cleaning with default options
normalizeColumnNames=true, columnNameStyle=snakeCase, trimStrings=true, collapseWhitespace=true, stringCase=preserve, treatConfiguredNullsAsNull=true, nullValues default, nullReplacement=null, no currency or European number conversion, no column or duplicate removal.