We provide IT Staff Augmentation Services!

Senior Data Engineer Resume

4.00/5 (Submit Your Rating)

ChicagO

SUMMARY

  • 8 years of extensive experience in Data Engineering using Apache Hadoop, HDFS, MapReduce, Pig, and Hive. Related skillset includes configuring & utilizingTalend, Protegrity, MAP R Streams,HBase, Zookeeper, Sqoop, Flume, OOZIE, Core Java, JavaScript, J2EE, Oracle 11G/10G, and HP PPM 9.2.
  • Experienced in data management & analysis usingHiveQL, Pig Latin, HBase, andcustom MapReduceprograms in Java.
  • Experienced in ingesting and exporting data using Sqoop from HDFS to RDBMS and vice - versa.
  • Experienced in NoSQL databases such as HBase and Cassandra.
  • Experienced in job workflow scheduling tools like Oozie and in managing Hadoop clusters using Cloudera Manager Tool.
  • Supported in setting up QA environments and updating configurations for implementing scripts with Pig and Sqoop.
  • Design and Develop ETL Processes in AWS Glue to migrate Campaign data from external sources like S3, ORC/Parquet/Text Files into AWS Redshift.
  • Worked on developing ETL pipelines on S3 parquet files on data late using AWS GLUE
  • Experiednce with AWS Cloud services like EC2, S3, EMR, RDS, Anthena and Glue.
  • Data Extraction, aggregations and consolidation of Adobe data within AWS Glue using PySpark.
  • Experience working with various Relational Databases (RDBMS): Oracle, MySQL, SQL Server, Complex Flat Files, Datasets, XML, and Flat files.
  • Used SSIS to create ETL packages to Validate, Extract, Transform and Load data into Data Warehouse and Data Mart.
  • Maintained and developed complex SQL queries, stored procedures, views, functions and reports that meet customer requirements using Microsoft SQL Server 2008 R2.
  • Created Views and Table-valued Functions, Common Table Expression (CTE), joins, complex subqueries to provide the reporting solutions.
  • Optimized the performance of queries with modification in T-SQL queries, removed the unnecessary columns and redundant data, normalized tables, established joins and created index.
  • Created SSIS packages using Pivot Transformation, Fuzzy Lookup, Derived Columns, Condition Split, Aggregate, Execute SQL Task, Data Flow Task and Execute Package Task.
  • Having good experience in writing and tuning SQL queries in Mongo DB, Teradata, Oracle, and SQL Server
  • 6+ yearsof experience in AWS, configuration, and deployment of workloads in AWS clonk
  • Created alerts to notify system outages or reaching threshold values. These alerts include Splunk license threshold limit, syslog server threshold limit, file system overflow and cold storage outage.
  • Developing customized Shell scripts in order to install, manage, configure multiple instances of Splunk forwarders, indexers, search heads, deployment servers.
  • Involved in Tuning SQL queries using Explain Plan and Tkprof utility to improve performance of the queries.
  • Involved in designing Technical specification document, identified various data mappings and the design of Data warehouse.
  • Loaded data using ETL tools like SQL loader and external tables to load data from data warehouse and various other database like SQL Server and DB2.
  • Developed tables with partitions and created indexes, constraints, triggers, synonyms, database links, table spaces, roles etc in different schema.
  • Created scripts to load the data into staging tables from other database using DBlinks.
  • Worked with various complex queries with joins, sub-queries, and nested queries in SQL queries.
  • Created detailed documentation for all the reports, alerts and dashboards. People without Splunk knowledge can follow the mentioned instructions and generate alerts/reports manually in case of automated mail generation failure (Firewall issues between SPLUNK and mail server).
  • Experience with Splunk User Interface in creating Dashboards and Visualizations.
  • Used SPLUNK forwarders to provide reliable and secure collection and delivery of data to the Splunk platform for indexing, storage and analysis.
  • Provide regular support guidance to Splunk project teams on complex solution and issue resolution.
  • Assisted administrators to ensure whether SPLUNK is actively and accurately running and monitoring on the current infrastructure implementation.
  • Development experience with PowerBuilder applications, application integration with relational database management systems, and experience working with other mid-tier programming languages, including Oracle PL/SQL, ADA, and C languages.
  • Designed and managed cloud infrastructures using Amazon Web Services (AWS), including VPC, EC2, S3, Elastic File System, RDS, Direct Connect, Route53, Cloud Watch, Cloud Trail, and Cloud Formation which allowed rapid prototyping and proof of concepts.
  • Configured and managed Elastic Load Balancing (ELB) to avoid single point of failure of applications, thus providing high availability and network load balancing.
  • Monitored resources and applications using AWS Cloud Watch, created alarms to key metrics for components likeEC2, ELB, RDS, S3, SNS, and configured notifications for the alarms generated based on events defined.
  • Worked on UNIX shell scripts usingK-shellfor the Scheduling sessions, automation of processes, pre and post-session scripts.
  • Experienced in agile methodology. Led SCRUM ceremonies.

TECHNICAL SKILLS

DWH / BI Tools: SQL Server services (SSIS & SSRS, SSAS), Business Intelligence Development Studio (BIDS), Visual Studio, Crystal Reports.

Hadoop Eco System: HDFS, MapReduce, MR Unit, YARN, Hive, Pig, HBase, Impala, Zookeeper, Sqoop, Oozie, DataStax Apache Cassandra, Flume,Spark

Tools: and Utilities: Talend, SQL Server Management Studio, SQL Server Enterprise Manager, SQL Server Profile, Visual Studio .Net, Microsoft Management Console, Business Intelligence Development Studio (BIDS), Crystal Reports, Microsoft Office

Languages: T-SQL, PL/SQL, Dynamic SQL, MDX, C, C++, C#, XML, Unix, Linux, Shell Scripting, Microsoft Azure, Java, JSP, Struts, JQuery, Javascript.

Databases: Oracle, MySQL, MS SQL Server 2014, 2012,2008R2,2008,2005,2000,Netezza, presto, teradata

Platforms/Operating Environments: Windows, Unix/Linux, AWS Cloud

Cloud Technologies: AWS EC2, EMR, AWS LAMBDA, S3, IAM, AWS GLUE, Snowflake, SnowSQL, SnowpipeAWS.

PROFESSIONAL EXPERIENCE

Confidential, Chicago

Senior Data Engineer

Responsibilities:

  • Imported and exported data from different databases, namely MySQL, PostgreSQL, Oracle, into HDFS and Hive using Sqoop for various migrations, PoCs.
  • Developed multiple MapReduce jobs in PIG and Hive for data cleaning and pre-processing.
  • Worked in setting up TWS jobs using Talend open studio, built Jenkins GitHub CI/CD pipelines for Supported projects.
  • Created ad-hoc reports to users in Tableau by connecting various data sources.
  • Preparing dashboards using calculated fields, parameters, calculations, groups, sets and hierarchies in Tableau.
  • Strong experience in migrating other databases to Snowflake.
  • Work with domain experts, engineers, and other data scientists to develop, implement, and improve upon existing systems.
  • Experience in analyzing data using HiveQL
  • Participate in design meetings for creation of the Data Model and provide guidance on best data architecture practices
  • Experience with Snowflake Multi - Cluster Warehouses.
  • Experience in Splunk reporting system.
  • Understanding of SnowFlake cloud technology.
  • Experience with Snowflake cloud data warehouse and AWS S3 bucket for integrating data from multiple source system which include loading nested JSON formatted data into snowflake table.
  • Professional knowledge of AWS Redshift
  • Experience in building Snowpipe.
  • Experience in using Snowflake Clone and Time Travel.
  • Experience in various data ingestion patterns to hadoop.
  • Participates in the development improvement and maintenance of snowflake database applications
  • Analyzed the source data and handled efficiently by modifying the data types. Used excel sheet, flat files, CSV files to generated Tableau ad-hoc reports.
  • Generated tableau dashboards with combination charts for clear understanding.
  • Created and tuned SQL queries in Mongo DB, Teradata, Oracle, and SQL Servers.
  • Worked on Hadoop technologies such as Splunk, Pig, Protegrity, Map R Streams.
  • Configured Zookeeper, Cassandra & Flume to the existing Hadoop cluster.
  • Converted Hive or SQL queries into Spark transformations using Python and Scala.
  • Optimized existing algorithms in Hadoop usingSparkContext,Spark-SQL,PySpark, Pair RDD's, andSpark YARN.
  • Developed Spark scripts by using Scala shell commands as per the requirement.
  • Utilized Spark API over Cloudera Hadoop YARN to perform analytics on data in Hive.
  • Developed Scala scripts, UDFFs using both Data frames/SQL/Data sets and RDD/MapReduce in Spark 1.6 for Data Aggregation, queries, and writing data back into OLTP system through Sqoop.
  • Improved the performance of Spark Applications for setting the right Batch Interval time, the correct level of Parallelism, and memory tuning.
  • Optimized existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames, and Pair RDD's.
  • Handled large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations, and others during the ingestion process itself.
  • Explored Kafka for the proof of concept for carrying out log processing on a distributed system.
  • Used Zookeeper and OOZIE Operational Services for coordinating the cluster and scheduling workflows.
  • Used max JVM parameters and Cursor Size inTalend as a part of Performance tuning
  • Explored many components in the palette to design Jobs & used Context Variables to Parameterize Talend Jobs.
  • Worked with connectors for Processing to improve job performance while working with the bulk data sources in the Talend
  • Skillfull inDBMSconcepts.
  • Built Dimensional Modelingand Data Models withStar, Snowflake, 3NF schemas forthe OLAPandODSapplications. designed and implementedETL Architecture
  • Tuned ETL processes & the SQL Queries for great performance.
  • Made complexdata analysisand provided crucial reports to support different departments.
  • Hands-on with the Data Visualization tools likeTableau and Business Intelligence tools likeBusiness Objects.
  • Generated Python Django forms to record data of online users and used PyTest for writing test cases.
  • Implemented and modified various SQL queries and Functions, Cursors and Triggers as per the client requirements.
  • Clean data and processed third party spending data into maneuverable deliverables within specific format with Excel macros and python libraries such as NumPy, SQLAlchemy and matplotlib.
  • Used Pandas as API to put the data as time series and tabular format for manipulation and retrieval of data.
  • Helped with the migration from the old server to Jira database (Matching Fields) with Python scripts for transferring and verifying the information.
  • Analyze Format data using Machine Learning algorithm by Python Scikit-Learn.
  • Used various versions of Hive, Spark, and Presto on multiple projects. Apart from regular queries, I have also implemented UDFs.
  • Experience in python, Jupyter, Scientific computing stack (numpy, scipy, pandasand matplotlib).
  • Perform troubleshooting, fixed and deployed many Python bug fixes of the two main applications that were a main source of data for both customers and internal customer service team.
  • Thoroughexperience in Python /Shell scriptingexperience for Process Automation and Scheduling.
  • Good exposure to Production support, Documentation, Development, Testing and Implementation.
  • Effective working relationships with client teams to understand and support requirements, strategic plans and develop tactical implement technology solutions, and pro-actively manage client expectations.

Environment: Hadoop, MapReduce, HDFS, Hive, HBase, Pig, Impala, Cassandra, Spark,Java, SQL, Tableau, PIG, Zookeeper, Sqoop, Teradata, presto, spark, Flume, Oozie, Redhat Linux.

Confidential, California

Data Engineer

Responsibilities:

  • Designed and built a data pipeline that streams data from client apps using web-sockets to the server and Kafka Consumer, which consumes that data and writes to HDFS data store. Different spark jobs from the HDFS store are reading this data using Spark-SQL and processing this data in stream and batch jobs.
  • Was responsible for creating on-demand tables on S3 files using Lambda Functions and AWS Glue using Python and PySpark.
  • Coordinated with team and Developed framework to generate Daily adhoc, Report’s and Extracts from enterprise data and automated using Oozie.
  • Worked on cloud deployments using maven, docker and Jenkins.
  • Designed and Co-ordinated with Data Science team in implementing Advanced Analytical Models in Hadoop Cluster over large Datasets.
  • Created monitors, alarms, notifications and logs for Lambda functions, Glue Jobs, EC2 hosts using Cloudwatch
  • Used AWS Glue for the data transformation, validate and data cleansing.
  • Used python Boto 3 to configure the services AWS glue, EC2, S3
  • Used AWS glue catalog with crawler to get the data from S3 and perform sql query operations
  • Worked on analyzing Hadoop clusters and different big data analytic tools, including Pig, Hive, and Sqoop.
  • Imported data from various data sources such as Oracle and Comptel server into HDFS using Sqoop and MapReduce transformations.
  • Analyzed the data by performing Hive queries and running Pig scripts to know user behavior like frequency of calls, top calling customers.
  • Developed mappings /Transformation/Job lets and designed ETL Jobs/Packages using Talend Integration Suite (TIS) in Talend 6.1
  • UsedTalend job let and various commonly usedTalend transformations components like tMap, tDie, tConvertType, tFlowMeter, tLogCatcher, tRowGenerator, tSetGlobalVar, tHashInput & tHashOutput and many more.
  • Responsible to Configure theHadoop cluster and troubleshoot common Cluster problems.
  • Cluster configuration and data transfer (distcp&hftp), inter and intra-cluster data transfer.
  • End-to-end deployment ownership for projects on AWS. This includes Python scripting for automation, scalability, builds promotions for staging to production, etc. Built a Continuous Integration environment using Jenkins, Nexus, Yum, and puppet.
  • Evaluated business requirements and prepared detailed specifications that follow project guidelines required to develop written programs.
  • Worked on Cloudera distribution and deployed on AWS EC2 Instances.
  • Hands on experience on Cloudera Hue to import data on to the graphical User Interface.
  • Responsible for building scalable distributed data solutions using Hadoop.
  • Involved in the loading of structured and unstructured data into HDFS.
  • Imported metadata from Relational Databases like Oracle, Mysql using Sqoop.
  • Implemented Web Interfacing with Hive and stored the data in Hive tables.
  • Loaded data from MySQL, a relational database to HDFS on regular basis using Sqoop Import/Export.
  • Responsible for implementing Map Reduce programs into Spark transformations using Spark and Scala .
  • Getting real time data using Kafka and processing using Spark and Scala.
  • Worked on loading CSV/TXT/AVRO/PARQUET files using Scala/Java language in Spark Framework and process the data by creating Spark Data frame and RDD and save the file in parquet format in HDFS to load into fact table using ORC Reader.
  • Configured and maintained the monitoring and alerting of production and corporate servers/storage using AWS Cloud Watch.
  • Designed and deployed multiple applications utilizing the AWS stack (EC2, Route53, S3, RDS, SNS, SQS, IAM), focusing on high availability, fault tolerance.

Environment: Hadoop, MapReduce, Spark, Pig, Hive, HDFS, Yarn, Hue, Oozie, Zookeeper, Impala, cluster health, Puppet, Flume, Sqoop, Kafka, KMS

Confidential, Maryland

Data Analyst

Responsibilities:

  • Experienced in all phases of Software development life cycle (SDLC) including System Analysis, Design, Data Modeling, Implementation and Support and maintenance of various applications in both OLTP and OLAP systems.
  • Possess strong Documentation skill and knowledge sharing among Team, conducted data modeling sessions for different user groups, facilitated common data models between different applications, participated in requirement sessions to identify logical entities
  • Extensive Experience working with business users as well as senior management.
  • Strong understanding of the principles of Data warehousing, Fact Tables, Dimension Tables, Star and Snowflake schema modeling.
  • Strong experience with Database performance tuning and optimization, query optimization, index tuning, caching and buffer tuning.
  • Extensive experience in relational Data modeling, Dimensional data modeling, logical/Physical Design, ER Diagrams and OLTP and OLAP System Study and Analysis.
  • Extensive experience in Enterprise Information Management and Architecture technologies including Information Lifecycle, Master Data Management and Business Intelligence.
  • Extensively used ERWIN and PowerDesigner to design Logical and Physical Data Models, to forward and reverse engineering data models and publishing data model to acrobat PDF files.
  • Excellent leadership, organization skills and requirement gathering skills such as JAD.
  • Strong experience in Database such as Microsoft Access, Oracle 11g/10g/9i/8i, SQL Server, DB2, Teradata, on Windows 2000/NT/98/95 platforms and latest.
  • Worked across all phases of the development life cycle from analysis through implementation and support.
  • Worked on supporting tasks that are included but not limited to code deployment, managing source control systems, virtual servers, scripting, etc.
  • Analyzed the system requirements concerning the business needs, created shell scripts, modified and executed validation scripts in multiple platforms.
  • Created, modified, and configured plans in Bamboo, integrated with the SVN repository, to compile and build EAR files that are essential for the deployment of applications.
  • Performed troubleshooting of ThinApp's
  • Maintained all the applications, websites, and components as part of managed services.

Environment: Agile Methodology,UNIX, MySQL, Oracle, C++, Java, GIT, SVN, VMWare View Client, Putty, RequestIt, JIRA, Artifactory, Bamboo, OD Tools, Atlassian Tools

We'd love your feedback!