Data Engineer Resume
Cincinnati, OH
SUMMARY
- 5+ years of experience in Data Engineer, including profound expertise and experience on statistical data analysis such as transforming business requirements into analytical models, designing algorithms, and strategic solutions that scales across massive volumes of data.
- Excellent Experience in Designing, Developing, Documenting, Testing of ETL jobs and mappings in Server and Parallel jobs using Data Stage to populate tables in Data Warehouse and Data marts.
- Establishes and executes the Data Quality Governance Framework, which includes end - to-end process and data quality framework for assessing decisions that ensure the suitability of data for its intended purpose.
- Expert in designing Server jobs using various types of stages like Sequential file, ODBC, Hashed file, Aggregator, Transformer, Sort, Link Partitioner and Link Collector.
- Expert in designing Parallel jobs using various stages like Join, Merge, Lookup, remove duplicates, Filter, Dataset, Lookup file set, Complex flat file, Modify, Aggregator, XML.
- Extensive experience in Text Analytics, generating data visualizations using R, Python and creating dashboards using tools like Tableau.
- Excellent knowledge of studying the data dependencies using metadata stored in the repository and prepared batches for the existing sessions to facilitate scheduling of multiple sessions.
- Utilized analytical applications like SPSS, Rattle and Python to identify trends and relationships between different pieces of data, draw appropriate conclusions and translate analytical findings into risk management and marketing strategies that drive value.
- Skilled in performing data parsing, data manipulation and data preparation with methods including describe data contents.
- Hands on experience with big data tools like Hadoop, Spark, Hive, Pig, Impala, Pyspark, Spark SQL.
- Good knowledge in Database Creation and maintenance of physical data models with Oracle, Teradata, Netezza, DB2, MongoDB, HBase and SQL Server databases.
- Experienced in writing complex SQL Quires like Stored Procedures, triggers, joints, and Sub quires.
- Interpret problems and provides solutions to business problems using data analysis, data mining, optimization tools, and machine learning techniques and statistics.
- Experience with Data Analytics, Data Reporting, Ad-hoc Reporting, Graphs, Scales, PivotTables and OLAP reporting.
- Ability to work with managers and executives to understand the business objectives and deliver as per the business needs and a firm believer in teamwork.
- Experience and domain knowledge in various industries such as healthcare, insurance, retail, banking, media and technology. Moreover, working closely with customers, cross-functional teams, research scientists, software developers, and business teams in an Agile/Scrum work environment to drive data model implementations and algorithms into practice.
- Strong written and oral communication skills for giving presentations to non-technical stakeholders.
TECHNICAL SKILLS
Databases: Oracle, MySQL, SQLite, NO SQL, RDBMS, SQL Server 2014, HBase 1.2, MongoDB 3.2.Teradata, Netezza. Cassandra. Alation, Data Governance
Database Tools: PL/SQL Developer, Toad, SQL Loader, Erwin.
Web Programming: Html, CSS, Xml, JavaScript.
Programming Languages: R, Python, SQL, Scala, UNIX, C, JAVA, Tableau
DWH BI Tools: Data Stage 9.1, 11.5, Tableau Desktop
Data Visualization: Dataiku, Tableau9.4/9.2.
Bigdata Framework: HDFS, MapReduce, Pig, Hive, Sqoop, Oozie, Zookeeper, Flume and HBase, Amazon EC2, S3 and Red Shift, Spark, Storm, Impala, Kafka.
Scheduling Tools: Autosys, Control-M.
Operating Systems: AIX, LINUX, UNIX.
Environment: AWS, AZURE, Databricks.com
Reporting Tools: SSIS/SSRS/SSAS.
PROFESSIONAL EXPERIENCE
Data Engineer
Confidential, Cincinnati, OH
Responsibilities:
- Work in a fast-paced agile development environment to quickly analyze, develop, and test potential use cases for the business.
- The individual will be responsible for design and development of High-performance data architectures which support data warehousing, real-time ETL and batch big-data processing.
- This project was mainly focus on reporting the commercial loan detailed information to ‘Federal department’ with applying ‘Data Governance controls on it.
- Working experience in Financial Reports “2052A”, “FR-Y9C”, “14Q”,10K-Q”
- Involved as primary on-site ETL Developer during the analysis, planning, design, development, and implementation stages of projects using IBM Web Sphere software (Quality Stage v9.1, Web Service, Information Analyzer, Profile Stage)
- Prepared Data Mapping Documents and Design the ETL jobs based on the DMD with required Tables in the Dev Environment.
- Active participation in decision making and QA meetings and regularly interacted with the Business Analysts &development team to gain a better understanding of the Business Process, Requirements & Design.
- Used DataStage as an ETL tool to extract data from sources systems, loaded the data into theORACLEdatabase.
- Designed and Developed Data stage Jobs to Extract data from heterogeneous sources, applied transform logics to extracted data and Loaded into Data Warehouse Databases.
- Created DataStage jobs using different stages like Transformer, Aggregator, Sort, Join, Merge, Lookup, Data Set, Funnel, Remove Duplicates, Copy, Modify, Filter, Change Data Capture, Change Apply, Sample, Surrogate Key, Column Generator, Row Generator, Etc.
- Experienced in developing parallel jobs using various Development/debug stages (Peek stage, Head & Tail Stage, Row generator stage, Column generator stage, Sample Stage) and processing stages (Aggregator, Change Capture, Change Apply, Filter, Sort & Merge, Funnel, Remove Duplicate Stage)
- Extensively worked with Join, Look up (Normal and Sparse) and Merge stages.
- Applying the Data modelling and Data Designing in-between staging and target for creating the views.
- Responsible for uploading the data into ‘Enterprise Data Warehouse’ by using the ETL tool ‘IBM DataStage 9.1 version or 11.5 version.
- Extensive knowledge on applying Data Governance controls.
- Working with RSR team and accomplish the tasks in timely manner and report to the stakeholders by weekly especially with the presentations.
- Experience with Data Analytics, Data Reporting, Ad-hoc Reporting, Graphs, Scales, PivotTables and OLTP reporting.
- Extensive knowledge on AQT tool. Classifying the domains in Enterprise Data Warehouse.
- Hands on experience on ‘DataIku’ visualization tool.
Environment: IBM Info sphere DataStage 9.1/11.5, Oracle 11g, Flat files, Autosys, UNIX, Erwin, TOAD, MS SQL Server database, XML files, MS Access database.
Data Engineer
Confidential, Bradenton, FL
Responsibilities:
- This project was focused on customer clustering. Used the ETL Data Stage Director to schedule and running the jobs, testing and debugging its components & monitoring performance statistics.
- Collaborated with EDW team in,High Level designdocuments for extract, transform, validate and load ETL process data dictionaries, Metadata descriptions, file layouts and flow diagrams.
- Develop an Estimation model for various product & services bundled offering to optimize and predict the gross margin
- A highly immersive Data Science program involving Data Manipulation & Visualization, Web Scraping, Machine Learning, Python programming, SQL, GIT, Unix Commands, NoSQL, MongoDB, Hadoop.
- Involved in creating UNIX shell scriptsfordatabase connectivity and executing queries in parallel job execution.
- Used the ETL Data Stage Director to schedule and running the jobs, testing and debugging its components & monitoring performance statistics.
- Performed scoring and financial forecasting for collection priorities using Python, and SAS.
- Handled importing data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS
- Managed existing team members lead the recruiting and on boarding of a larger Data Science team that addresses analytical knowledge requirements.
- Developed predictive causal model using annual failure rate and standard cost basis for the new bundled services.
- Design and develop analytics, machine learning models, and visualizations that drive performance and provide insights, from prototyping to production deployment and product recommendation and allocation planning.
- Worked with sales and marketing team for Partner and collaborate with a cross-functional team to frame and answer importantdataquestions.
- Prototyping and experimenting ML algorithms and integrating into production system for different business needs.
- Successfully implemented pipeline and partitioning parallelism techniques and ensured load balancing of data.
- Deployed different partitioning methods like Hash by column, Round Robin, Entire, Modulus, and Range for bulk data loading and for performance boost.
- Worked on Multiple datasets containing 2billion values which are structured and unstructureddata about web applications usage and online customer surveys
- Design built and deployed a set of python modeling APIs for customer analytics, which integrate multiple machine learning techniques for various user behavior prediction and support multiple marketing segmentation programs.
- Used classification techniques including Random Forest and Logistic Regression to quantify the likelihood of each user referring.
- Designed and implemented end-to-end systems forDataAnalytics and Automation, integrating custom visualization tools using R, Tableau, and Power BI.
Environment: IBM DataStage, Python, Spark framework, Redshift, MS Excel, NoSQL, Tableau, T-SQL, ETL, RNN, LSTM MS Access, XML, MS office 2007, Outlook, MS SQL Server.
Data Engineer
Confidential, Sanjose, CA
Responsibilities:
- Responsible for analyzing large data sets to develop multiple custom models and algorithms to drive innovative business solutions.
- Perform Data profiling, preliminary data analysis and handle anomalies such as missing, duplicates, outliers, and imputed irrelevant data. Remove outliers using Proximity Distance and Density based techniques.
- Involved in Analysis, Design and Implementation/translation of Business User requirements.
- Used supervised, unsupervised and regression techniques in building models.
- Performed Market Basket Analysis to identify the groups of assets moving together and recommended the client their risks
- Determined trends and significant data relationships using advanced Statistical Methods.
- Implemented techniques like forward selection, backward elimination and step wise approach for selection of most significant independent variables.
- Performed Feature selection and Feature extraction dimensionality reduction methods to figure out significant variables.
- Used RMSE score, Confusion matrix, ROC, Cross validation and A/B testing to evaluate model performance in both simulated environment and real world.
- Performed Exploratory Data Analysis using R. Also involved in generating various graphs and charts for analyzing the data using Python Libraries.
- Involved in the execution of multiple business plans and projects Ensures business needs are being met Interpret data to identify trends to go across future data sets.
- Developed interactive dashboards, created various Ad Hoc reports for users in Tableau by connecting various data sources.
Environment: Python, SQL server, Hadoop, HDFS, HBase, MapReduce, Hive, Impala, Pig, Sqoop, Mahout, LSTM, RNN, Spark MLLib, MongoDB, Tableau, Unix/Linux.
Data Analyst
Confidential, Phoenix, AZ
Responsibilities:
- Involved in Analysis, Design and Implementation/translation of Business User requirements.
- Worked on collection of large sets using Python scripting. Spark SQL
- Worked on large sets of Structured and Unstructured data.
- Worked on creating DL algorithms using LSTM and RNN.
- Actively involved in designing and developing data ingestion, aggregation, and integration in Hadoop environment.
- Developed Sqoop scripts to import export data from relational sources and handled incremental loading on the customer, transaction data by date.
- Experience in creating Hive Tables, Partitioning and Bucketing.
- Performed data analysis and data profiling using complex SQL queries on various sources systems including Oracle 10g/11g and SQL Server 2012.
- Identified inconsistencies in data collected from different source.
- Worked with business owners/stakeholders to assess Risk impact, provided solution to business owners.
- Experienced in determine trends and significant data relationships Analyzing using advanced Statistical Methods.
- Carrying out specified data processing and statistical techniques such as sampling techniques, estimation, hypothesis testing, time series, correlation and regression analysis Using R.
- Applied various data mining techniques: Linear Regression & Logistic Regression, classification, clustering.
- Took personal responsibility for meeting deadlines and delivering high quality work.
- Strived to continually improve existing methodologies, processes, and deliverable templates.
Environment: R, SQL server, Oracle, HDFS, HBase, MapReduce, Hive, Impala, Pig, Sqoop, NoSQL, Tableau, RNN, LSTM, Unix/Linux, Core Java.
Data Analyst
Confidential
Responsibilities:
- Worked on different dataflow and control flow task, for loop container, sequence container, script task, executes SQL task and Package configuration.
- Created new procedures to handle complex logic for business and modified already existing stored procedures, functions, views and tables for new enhancements of the project and to resolve the existing defects.
- Loading data from various sources like OLEDB, flat files to SQL Server 2012 database Using SSIS Packages and created data mappings to load the data from source to destination.
- Created batch jobs and configuration files to create automated process using SSIS.
- Created SSIS packages to pull data from SQL Server and exported to Excel Spreadsheets and vice versa.
- Built SSIS packages, to fetch file from remote location like FTP and SFTP, decrypt it, transform it, mart it to data warehouse and provide proper error handling and alerting
- Extensive use of Expressions, Variables, Row Count in SSIS packages
- Data validation and cleansing of staged input records was performed before loading into Data Warehouse
- Automated the process of extracting the various files like flat/excel files from various sources like FTP and SFTP (Secure FTP).
- Deploying and scheduling reports using SSRS to generate daily, weekly, monthly and quarterly reports.
Environment: MS SQL Server 2005 & 2008, SQL Server Business Intelligence Development Studio, SSIS-2008, SSRS-2008, Report Builder, Office, Excel, Flat Files, .NET, T-SQL.
