Big Data Engineer Resume
Cincinnati, OH
SUMMARY
- Around 7 years of expertise in Enterprise Application Development, Web Applications, Client - Server Technologies, Java Web Programming, and Big Data technologies design and deployment.
- Hadoop architecture and ecosystem expertise, including HDFS, MapReduce, Pig, Hive, Sqoop Flume, and Oozie.
- Thorough understanding of Hadoop ecosystem tools such as HDFS, MapReduce, YARN, Spark, Kafka, Hive, Sqoop, Pig, Impala, HBase, Flume, Oozie, and Zookeeper.
- Solid understanding of Hadoop architecture, as well as hands-on expertise with Hadoop components such as Job Tracker, Task Tracker, Name Node, Data Node, and Map Reduce principles, as well as the HDFS Framework.
- Experience in monitoring Hadoop clusters using Cloudera Manager and Web UI.
- Experience in implementation of various Machine Learning models and have good understanding of concepts behind it.
- Extensive familiarity with Hive and Spark User Defined Functions (UDFs). Using Apache Sqoop, I was able to import and export data from HDFS to RDBMS/NoSQL databases and vice versa.
- Worked with HBase, Cassandra, Mongo DB, Azure Cosmos DB and Spark-Redis are examples of NoSQL databases.
- Have developed Spark applications using Spark - SQL in Databricks for data extraction, transformation and aggregation from numerous file formats, as well as analyzing and manipulating data to reveal insights regarding client usage trends.
- Knowledge of automated deployments using Azure Resource Manager Templates, DevOps, and Git repositories for automation, as well as Continuous Integration and Continuous Delivery (CI/CD).
- Expertise in using Oozie and Airflow to schedule workflows and create Job Designers. Extensive expertise of Amazon AWS technologies such as EMR and EC2 for processing massive amounts of data quickly and efficiently, as well as Machine Learning models.
- Used Azure Data platform capabilities such as Azure Data Lake, Azure Data Factory, HDInsight, Azure SQL Server, Azure Machine Learning and Power BI to build huge Lambda systems.
- Experience in working with AWS Cloud services and SDKs to work with services like AWS API Gateway, Lambda, S3, IAM, and EC2.
- Using Linked Services/Datasets/Pipeline/, created Pipelines in ADF to extract, transform, and load data from various sources such as Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool, and backwards.
- Experience in building, deploying, troubleshooting, and data extraction for huge number of records using Azure Data Factory (ADF).
- Good Knowledge on Azure Synapse Analytics architecture and component integrations
- Proactively monitored Synapse workload (Dedicated SQL pool and Spark pool) and engaged internal / external teams as appropriate to resolve errors / failures / throughput bottleneck
- Configured and deployed EC2, VPC, S3, IAM, ELB, AWS Lambda, DynamoDB, CloudWatch, CloudFormation, AWS Auto Scaling, and Route53, among other AWS services.
- Good knowledge on Core Java, J2EE, JDBC, Servlets, JSP, Exception Handling, Multithreading, EJB, XML, HTML5, CSS3, JavaScript, and AngularJS.
- Experience in Google Cloud components, Google container builders and GCP client libraries and cloud SDK’s.
- Worked with a variety of databases, including Oracle, SQL Server, and MySQL, as well as written stored procedures, functions, joins, and triggers for various Data Models.
- Experience in building and architecting multiple Data pipelines, end to end ETL and ELT process for Data ingestion and transformation in GCP.
- Creating and monitoring batches and sessions using informatica Power center server.
- Analyzed the data sources and targets using Informatica data profiling option.
- Design and develop ETL processes. Establish coding standards, perform code review and automate the health of the platform
- Perform data analysis, develop database design, share, and get consensus on design, and then develop the stored procedures to fulfill the database design
- Manages end-user accounts and accessibility; provides technical expertise to end-users who create complex queries and reports
- Experience with the Python libraries NumPy, Matplotlib, Pandas, Seaborn, and Plotly. Used PySpark, NumPy, and pandas on big datasets.
- Extensive knowledge of relational and non-relational databases, including Oracle, PL/SQL, SQL Server, MySQL, and DB2.
- Scripts and batch jobs were created to schedule various Hadoop programs. TensorFlow was used to train the model from useful data.
- Using Apache Kafka and Flume, we flooded HDFS with massive amounts of data.
- Installation and integration of Kafka with Spark Streaming knowledge.
- Experience with Snowflake Multi-Cluster Warehouses, Snowflake Databases, Schemas, and Table Structures.
- Able to handle numerous priorities and workloads, as well as quickly learn and adapt to new technologies and environments.
- Experience in using Databricks in creation of notebooks and scheduling jobs.
TECHNICAL SKILLS
Hadoop Services: DFS, Map Reduce, Spark, YARN, Pig, Impala, Sqoop, Flume, Kafka, Stream-sets, NiFi, Hue
Hadoop Distribution: Cloudera, Hortonworks
Databases: HBase, Spark-Redis, Cassandra, Oracle, MySQL, Azure Cosmos, Mongo, Postgress, Teradata, Hive, Versant
Scheduling Tools: AutoSys, Oozie, Airflow
Cloud Computing Tools: AWS (Amazon EC2, Amazon EMR, LAMBDA, Amazon GLUE, Amazon S3, ATHENA), Azure (Azure Data Lake, Azure Data Factory, Azure Databricks, Azure SQL Database, Azure SQL data Warehouse), GCP.
Programming Languages: Python, Java, Scala, SQL, PL/SQL, Pig Latin, HiveQL, Unix, JavaScript, Shell Scripting.
Java & J2EE Technologies: Core Java, Servlets, Hibernate, Spring, Struts, JMS, EJB.
Operating Systems: UNIX, Windows, LINUX. Build Tools Jenkins, Maven, ANT.
Open-Source Software: Kubernetes, Docker, Git Hub, Rundeck
Development Tools: Eclipse, NetBeans, Microsoft SQL Studio, Toad.
PROFESSIONAL EXPERIENCE
Big Data Engineer
Confidential, Cincinnati, OH
Responsibilities:
- Implemented Hadoop framework to capture user navigation across the application to validate the user interface and provide analytic feedback/result to the UI team.
- Loaded data into the cluster from dynamically generated files using Flume and from relational database management systems using Sqoop.
- Involved in performance tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and memory tuning.
- Built a real time streaming pipeline utilizing Kafka, Spark Streaming and Redshift.
- Developed logical and physical data flow models for Informatica ETL applications.
- Worked on creation of customer Docker container images, tagging, and pushing of data images.
- Written Hive queries for data analysis to meet the business requirements.
- Implemented and analyzed SQL query performance issues in databases.
- Responsible for design development of Spark SQL Scripts based on Functional Requirements and Specifications.
- Responsible for implementing various ML models both Supervised and Unsupervised models.
- Hands on experience in loading data from UNIX file system to HDFS.
- Experienced on loading and transforming of large sets of structured, semi structured and unstructured data from HBase through Sqoop and placed in HDFS for further processing.
- Managing and scheduling Jobs on a Hadoop cluster using Oozie.
- Involved in creating Hive tables, loading data, and running hive queries in those data.
- Extensive Working knowledge of partitioned table, UDFs, performance tuning, compression-related properties, thrift server in Hive.
- Worked on Microsoft Azure services like HDInsight clusters, BLOB, ADLS, Data Factory and Logic Apps and done POC on Azure Data Bricks.
- Monitored Azure Synapse Analytics using Dynamic Management views to identify the performance bottlenecks.
- CreateCosmosscripts which pulls data from upstream structured streams and include business logic & transformations to meet the requirements.
- Worked with Parquet file formats to process batch and interactive workloads.
- Experience in developing spark applications using spark-SQL in databricks for data extraction, transformation, and aggregation from multiple file formats for analyzing and transforming the data to uncover insights into the customer usage patterns.
- Extract Transform and Load data from Sources Systems to Azure Data Storage services using a combination of Azure Data Factory, T-SQL, Spark SQL, and U-SQL Azure Data Lake Analytics. Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing the data in InAzure Databricks.
- Created Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform, and load data from different sources like Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards.
- Developed JSON Scripts for deploying the Pipeline in Azure Data Factory (ADF) that process the data using the SQL Activity.
- Used ELK cluster to organize error codes and log information for continuous monitoring spark applications and created alters in Kibana dashboard to trigger emails to production support team distribution list for any job failure.
- Troubleshooting any build issue with ELK and work towards the solution.
- Participated in problem resolving, change, release, and event management for ELK stack.
- Designed end to end scalable architecture to solve business problems using various Azure Components like HDInsight, Data Factory, Data Lake, Storage and Machine Learning Studio. Developed JSON Scripts for deploying the Pipeline in Azure Data Factory (ADF) that process the data using the SQL Activity.
- Worked alone on Google Cloud Migration project from Azure PaaS and converted their ETLs through Databricks and SSIS. Scheduled their jobs thru Databricks Jobs and SQL Server deployed packages with SQL Server Agent.
- Worked on Databricks in creation of multiple notebooks as per the requirements and scheduling jobs using Apache Airflow.
Environment: Azure, ADF, Azure Synapse, Apache Hadoop, MapReduce, HDFS, CentOS 6.4, Spark, Kafka, HBase, Hive, Pig, Oozie, Flume, Java (jdk 1.6), Eclipse, ELK.
Data Engineer
Confidential, Rockledge, FL
Responsibilities:
- Used Spark to import customer information data from Oracle database into HDFS for data processing along with minor cleansing.
- Developed MapReduce jobs to calculate the total usage of data by commercial routers in different locations using Horton work distribution.
- Involved in information gathering for new enhancements in Spark, Production support for field issues and label installs for Hive scripts and MapReduce jobs.
- Managed GIT repositories for branching, merging, and tagging.
- Used Maven to build rpms from source code in Scala checked out from GIT repository, with Jenkins being the Continuous Integration Server and Artifactory as repository manager.
- Responsible for Setting up UNIX/Linux environments for various applications using shell scripting.
- Set up scalability for application servers using command line interface for Setting up and administering DNS system in AWS using Route53.
- Managed users and groups using the Amazon Identity and Access Management (IAM).
- In-depth understanding of the principles and best practices ofSoftwareConfiguration
- Designed and developed server-side applications on Linux platform in a fast-paced environment.
- Translated customer business requirements into technical design documents, established specific solutions, and leading the efforts including programming in Spark Scala and testing that culminate in client acceptance of the results.
- Expertise in Object-Oriented Design (OOD) and end-to-endsoftwaredevelopment experience working on Scala coding and implementing mathematical models in Spark Analytics.
- Softwaretranslating the business requirements into Use Cases and Diagrams conducting reviews of Codes and test cases analyzing change requests enhancements managing release plans for business apps
- Developed oozie workflows and scheduling jobs through Hue.
- Used AWS Lambda to perform data validation, filtering, sorting and other transformations for every data change in a HBase table and load the transformed data to RDS.
- Loading data from different servers to S3 bucket and setting appropriate bucket permissions.
- Configured routing to send JMS files to interact with application for real time data using Kafka.
- Managed Zookeeper for cluster co-ordination and Kafka Offset monitoring.
- Optimized legacy queries to extract the customer information from Oracle.
- Reviewed HDFS usage and system design for future scalability and fault tolerance.
- Managed and reviewed Hadoop log files using Spark to identify issues when job fails.
Environment: HDFS, Spark, Scala, Python, Shell Scripting, Hive, HBase, Oracle, MapReduce, Logstash, Jenkins, Versant, Java, Kafka, Horton works, GIT, ClearCase, Zookeeper, Ansible, AWS.
Big Data Engineer
Confidential
Responsibilities:
- Responsible for building scalable distributed data solutions using Hadoop.
- Used Spark for improving performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark Sessions, Spark-SQL, Data Frame, Pair RDD's, Spark YARN.
- Designed and developed POCs in Spark using Scala to compare the performance of Spark with Map Reduce, Hive.
- Involved in creating Hive tables and loading and analyzing data using hive queries.
- Involved in design and development of an application using Hive (UDF).
- Used HUE for Hive Query execution.
- Used Hive QL to analyze the partitioned and bucketed data, Executed Hive queries on Parquet tables stored in Hive to perform data analysis to meet the business specification logic.
- Developed spark jobs using SCALA and PYTHON on top of YARN/MR v2 for interactive and batch analysis.
- Developed PYTHON script to start and end a job smoothly for a UC4 workflow.
- Built real timedata pipelinesby developingKafkaproducers andSparkstreaming applications for consuming.
- Written Hive queries on the analyzed data for aggregation and reporting.
- Implemented Spark Scripts using Scala, Spark SQL to access hive tables into spark for faster processing of data.
- Developed end to end data processing pipelines that begin with receiving data using distributed messaging systemsKafkathrough persistence of data intoHBase.
- Created a Python based web application using Python scripting for data processing, MySQL for the database, and HTML/CSS/JQuery and HighCharts, matplotlib for data visualization of sales, tracking progress, identifying trends.
- Developed various Python scripts to find vulnerabilities with SQL Queries by doing SQL injection, permission checks and performance analysis.
- Used Hive Context to integrate Hive metastore and Spark SQL for optimum performance.
- Developed Sqoop Jobs to load data from RDBMS to external systems like HDFS and HIVE.
- UsedInformaticaDesignertocreate,load,updatemappingsusingdifferenttransformationstomoveDataMartsinDataWarehouse.
- PerformeddataanalysisanddataprofilingusingSQLandInformaticaDataExploreronsourcessystemsincludingOracleandTeradata.
- AnalyzedtheETLprocesscreatedinInformatica,developedSQLqueriesandtableswhichreplic-atedtheETL,andcomparedtheSQLresultswiththeInformaticatables.
- WroteSQLscriptstotestmappingsanddevelopedTraceabilityMatrixtotestscriptsforanyChangeControlinrequirementsleadingtotestcaseupdate.
- ReviewedbasicSQLqueriesandeditedinner,left,andrightjoinsinTableauDesktopbyconnectinglive/dynamicandstaticdatasets.
- CreatedvisuallyimpactfuldashboardsinExcelandTableaufordatareportingwithrespecttocu-stomer,policyinformationandanalyzedtoidentifythekeymetrics.
- Launching Amazon EC2 Cloud Instances using Amazon Images (Linux/ Ubuntu) and configuring launched instances with respect to specific applications.
- Added support for AWS S3 and RDS to host static /media files and the database into amazon cloud.
- Ingested data from disparate data sources using a combination of SQL, Google Analytics API, and salesforce API using python to create data views to be used in BI tools like Tableau.
- Worked in Agile Iterative sessions to create Hadoop Data Lake for the client
- Developed Stored Procedures and Functions, Views for the Oracle database PL/SQL.
- Responsible for generating actionable insights from complex data to drive real business results for various application teams and worked in Agile Methodology projects extensively.
Environment: Amazon services, Spark SQL, PYTHON, CDH, HDFS, Hive, Pig, Apache Sqoop, Java (JDK SSE 6, 7), Scala, Shell scripting, Linux, MYSQL Oracle Enterprise DB, Eclipse, Oracle, Git, Oozie, Informatica, Tableau, MySQL, Soap, and Agile Methodologies.
Java Developer
Confidential
Responsibilities:
- Responsible for acquiring, assessing, and translating requirements into technical specifications.
- Created sequence and class diagrams.
- Created presentation layer with JSP, Java, HTML, and JavaScript.
- Used Spring Core Annotations for Dependency Injection.
- To enable dynamic fetching and displaying of diverse table data using JSF tag libraries, I designed and created a ‘Convention Based Coding’ using Hibernates persistence framework and O-R mapping capability.
- For database connectivity and accessing the session for database transactions, I designed and built Hibernate configuration and the session-per-request design pattern. For retrieving and saving data in databases, I used HQL and SQL.
- Contributed to the design and implementation of database schema and Entity-Relationship diagrams for the application's backend Oracle database tables.
- Used Apache Axis to implement web services.
- Designed and developed Oracle Stored Procedures and Triggers to meet the application's requirements.
- Created sophisticated SQL queries to pull data from the database SOAP web service interfaces implemented in Java were designed and built.
- For the build process, Apache Ant was used.
Environment: Java, JDK 1.5, Servlets, Hibernate, Ajax, Oracle 10g, Eclipse, Apache Ant, Web Services (SOAP), Apache Axis, Apache Ant, Web Logic Server, JavaScript, HTML, CSS, XML.
