We provide IT Staff Augmentation Services!

Sr. Data Engineer Resume

0/5 (Submit Your Rating)

Ada, MI

SUMMARY

  • Over 8+ years of progressive complex IT experience in Software Life Cycle Development including analysis, design (system/database/OO), development, deployment, testing, documentation, implementation & maintenance of application software in Web - based environments, and Client/Server architectures.
  • Expertise in AWS EMR and Spark deployment with S3 file system.
  • Working Experience on designing and implementing complete end-to-end Hadoop infrastructure using MapReduce, Hive, PIG, Sqoop, Oozie, Flume, Spark, HBase and Zookeeper.
  • Extensively worked on Spark using Python and Scala on cluster for computational (analytics), installed it on top of Hadoop.
  • Dealt with Improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, Pair RDD's, YARN.
  • Expert in PySpark, Python, Pysql and design technique as well as experience working across large environments with multiple operating systems.
  • Working experience in various Linux server environments from DEV all the way to PROD and along with cloud powered strategies embracing AWS and Azure.
  • Exposure to job workflow scheduling and monitoring tools like Oozie (hive, pig) and DAG (lambda).
  • Installed and configured Flume, Hive, Pig, Sqoop, HBase on the Hadoop cluster.
  • Hands on experience in Automation of Sqoop incremental imports by using Sqoop and automating jobs using Oozie.
  • Extensively used Stash, Bit-Bucket and GITHUB for the code control purpose.
  • Expertise in Tracking, Documenting, Capturing, Managing and Communicating the Requirements using Requirement Traceability Matrix (RTM)which helped in controlling numerous artifacts produced by the teams across the deliverables for a project.
  • Developed Critical Items project in which the change analysts will check for the critical items in the change order and will release the change order from Agile.
  • Extensive experience in using SQL and PL/SQL to write Stored Procedures, Functions, Packages, snapshots, Triggers, and optimization with Oracle, DB2 and MySQL databases.
  • Experienced in designing, built, and deploying a multitude of applications utilizing almost all the AWS stack (Including EC2, R53, S3, RDS, DynamoDB, SQS, IAM, and EMR), focusing on high availability, fault tolerance, and auto-scaling.

PROFESSIONAL EXPERIENCE

Sr. Data Engineer

Confidential, ADA, MI

Responsibilities:

  • Importing data from different sources like customer database, MySQL database, MongoDB, SFTP folder for converting raw data in structured format and analyzing the different patterns/trends of customers
  • Utilizing python/API calls to import data from databases like MySQL, PostgreSQL, MongoDB and different web applications
  • Importing data from various data sources with Apache Airflow and performing the data transformation using Hive, MySQL and loading data in HDFS
  • Parsing the data from S3 through the Python API calls through the Amazon API Gateway generating Batch Source for processing
  • Using Spark, performed various transformations/actions and the result data is saved back to HDFS from there to target database Snowflake
  • Migrating data from external sources like S3, used text/parquet file formats to be loaded into AWS redshift and designing, developing ETL processes in AWS Glue
  • Developing the code in PySpark to transform raw data into structured format jobs on EMR
  • Utilizing Apache Spark to write the code in PySpark and SparkSQL to process data on Amazon EMR and performing necessary transformation
  • Processing the data with stateless and state full transformations with Spark Streaming programs to process near real time data from Kafka and Kinesis
  • Building the structured data model with elastic search by using Python/Spark and developing the ETL pipelines for further analysis
  • Receiving event from S3 bucket by creating lambda deployment function and configuring it
  • Writing AWS Lambda functions in python for AWS's Lambda which is invoking python scripts to perform various transformations and analytics on large data sets in EMR clusters
  • Deploying AWS Lambda code from Amazon S3 buckets and implementing a 'serverless' architecture using API Gateway, Lambda, and Dynamo DB
  • Developing Airflow Workflow to schedule batch and real-time data from source to target
  • Monitor Resources and Applications using AWS Cloud Watch, including creating alarms to monitor metrics such as EBS, EC2, ELB, RDS, S3, SNS and configured notifications for the alarms generated based on events defined
  • Working on Apache Airflow for data Ingestion, Hive & Spark for data processing & Oozie for designing complex workflows in Hadoop framework

Environment: AWS (EC2, S3, EBS, ELB, RDS, Cloud Watch), Cassandra, PySpark, Apache Spark, HBase, Apache Kafka, Hive, Kinesis, Python, Spark streaming, Machine Learning, Snowflake, Oozie, Tableau, Power BI, NoSQL, PostgreSQL, Shell Scrip, Scala

Data Analytics Engineer

Confidential, Columbus, Indiana

Responsibilities:

  • Chaired a team of 2 people for telecom inventory maintenance of 5+ customer using SQL server and achieved historic annual savings of 8% with service providers AT&T, Verizon, Granite Telecom, CenturyLink.
  • Used SQL on data sets to provide ad-hoc data requests and created a report metrics which accelerated reporting by 3%
  • Developed SQL queries of 8% efficiency to fetch required data, analyzed expense based on requirements
  • Analysed the sql scripts and designed it by using PySpark SQL for faster performance.
  • Built Jupyter notebooks using PySpark for extensive data analysis and exploration.
  • Implemented code coverage and integrations using Sonar for improving code testability.
  • Pushed application logs and data streams logs to Applications Insights for monitoring and alerting purpose.
  • Worked on migrating data from HDFS to Azure HD Insights and Azure Databricks.
  • Experience designing solutions in Azure tools like Azure Data Factory, Azure Data Lake, SQL DWH, Azure SQL & Azure SQL Data Warehouse, Azure Functions.
  • Migrated existing processes and data from our on-premises SQL Server and other environments to Azure Data Lake.
  • Used Azure Databricks for fast, easy and collaborative spark-based platform on Azure.
  • Used Databricks to integrate easily with the whole Microsoft stack.
  • Developed spark applications in python (PySpark) on distributed environment to load huge number of CSV files with different schema in to Hive ORC tables.

Data Engineer

Confidential, Fort lauderdale, FL

Responsibilities:

  • Actively worked with Business Partners and Designers to get the Requirements and create High Level Application Design.
  • Preparing Dashboards using calculations, parameters in Tableau and Developing Tableau reports that provide clear visualizations of various industry specific KPIs.
  • Access and change extremely large datasets through filtering, grouping, aggregation, and statistical calculation.
  • Consult with customers and inside partners to accumulate necessities and set milestones in the research, advancement, and execution periods of the project lifecycle.
  • Understanding of Tableau features like calculated fields, parameters, table calculations, row-level security, R integration, joins, data blending, and dashboard actions.
  • Develop, organize, manage, and maintain graph, table, slide and document templates that will allow for efficient creation of reports.
  • Created cubes using packages which are built in framework manager.
  • Created and supported congas Transformer models based on the dimensions, levels and measures required for the analysis studio.
  • Modify the existing report based upon the change request by the client.
  • Developed complex reports using Drill through, conditional blocks and render variables.
  • Involved in status calls for requirement gathering and updating the requirements in the Design document and publishing the updated document into share point.
  • Involved in unit testing of reports and model.
  • Scheduling and Distributing Reports through Schedule Management in Congo’s Connection.

Environment: Tableau Server, Tableau Desktop, Teradata SQL Assistant, Teradata Administrator, SQL, Share Point, Agile-Scrum, Microsoft Office Suite.

Data Analyst/Engineer

Confidential, Philadelphia, PA

Responsibilities:

  • Responsible for gathering requirements from Business Analyst and Operational Analyst and identifying the data sources required for the request.
  • Created value from data and drive data-driven decisions by performing advanced analytics and statistical techniques to determine to deepen insights, optimal solution architecture, efficiency, maintainability, and scalability which make predictions and generate recommendations.
  • Enhancing data collection procedures to include information that is relevant for building analytic systems.
  • Worked closely with a data architect to review all the conceptual, logical and physical database design models with respect to functions, definition, maintenance review and support data analysis, Data quality and ETL design that feeds the logical data models.
  • Maintained and developed complex SQL queries, stored procedures, views, functions, and reports that qualify customer requirements using SQL Server 2012.
  • Creating automated anomaly detection systems and constant tracking of its performance.
  • Support Sales and Engagement's management planning and decision making on sales incentives.
  • Used statistical analysis, simulations, predictive modelling to analyze information and develop practical solutions to business problems.
  • Extending the company's data with third-party sources of information when needed.
  • Précised development of several types of sub-reports, drill down reports, summary reports, parameterized reports, and ad-hoc reports using SSRS through mailing server subscriptions &SharePoint server.

Environment: SQL Server 2012, SSRS, SSIS, SQL Profiler, Tableau, Qlik View, Agile, ETL, Anomaly detection.

ETL Developer

Confidential

Responsibilities:

  • Designed and built terabyte, full end-to-end Data Warehouse infrastructure from the ground up on Redshift for large scale data handling Millions of records
  • Developed workflows in Oozie for business requirements to extract the data using Sqoop
  • For data exploration stage used Hive to get important insights about the processed data from HDFS
  • Worked on Big data on AWS cloud services i.e., EC2, S3, EMR and DynamoDB
  • Expertise knowledge in Hive SQL, Presto SQL and Spark SQL for ETL jobs and using the right technology for the job to get done
  • Responsible for ETL and data validation using SQL Server Integration Services
  • Designed and Developed ETL jobs to extract data from Salesforce replica and load it in data mart in Redshift
  • Involved in Data Extraction from Oracle and Flat Files using SQL Loader Designed and developed mappings using Informatica
  • Developed PL/SQL procedures/packages to kick off the SQL Loader control files/procedures to load the data into Oracle
  • Build and maintain complex SQL queries for data analysis, data mining and data manipulation
  • Developed Matrix and tabular reports by grouping and sorting rows
  • Actively participated in weekly meetings with the technical teams to review the code
  • Participate in requirement gathering and analysis phase of the project in documenting the business requirements by conducting workshops/meetings with various business users
  • Built machine learning algorithms like linear regression, decision tree, random forest for continuous variable problems, estimated machine learning algorithm's performance for time series data
  • Analyzed large data sets using pandas to identify different trends/patterns about data
  • Developed schemas to handle reporting requirements using Tableau.

We'd love your feedback!