We provide IT Staff Augmentation Services!

Big Data Engineer Resume

0/5 (Submit Your Rating)

Cincinnati, OH

SUMMARY

  • Around 7 years of expertise in Enterprise Application Development, Web Applications, Client - Server Technologies, Java Web Programming, and Big Data technologies design and deployment.
  • Hadoop architecture and ecosystem expertise, including HDFS, MapReduce, Pig, Hive, Sqoop Flume, and Oozie.
  • Thorough understanding of Hadoop ecosystem tools such as HDFS, MapReduce, YARN, Spark, Kafka, Hive, Sqoop, Pig, Impala, HBase, Flume, Oozie, and Zookeeper.
  • Solid understanding of Hadoop architecture, as well as hands-on expertise wif Hadoop components such as Job Tracker, Task Tracker, Name Node, Data Node, and Map Reduce principals, as well as teh HDFS Framework.
  • Experience in monitoring Hadoop clusters using Cloudera Manager and Web UI.
  • Experience in implementation of various Machine Learning models and has good understanding of concepts behind it.
  • Extensive familiarity wif Hive and Spark User Defined Functions (UDFs). Using Apache Sqoop, I was able to import and export data from HDFS to RDBMS/NoSQL databases and vice versa.
  • Worked wif HBase, Cassandra, Mongo DB, Azure Cosmos DB and Spark-Redis are examples of NoSQL databases.
  • Has developed Spark applications using Spark - SQL in Databricks for data extraction, transformation and aggregation from numerous file formats, as well as analyzing and manipulating data to reveal insights regarding client usage trends.
  • Knowledge of automated deployments using Azure Resource Manager Templates, DevOps, and Git repositories for automation, as well as Continuous Integration and Continuous Delivery (CI/CD).
  • Expertise in using Oozie and Airflow to schedule workflows and create Job Designers. Extensive expertise of Amazon AWS technologies such as EMR and EC2 for processing massive amounts of data quickly and efficiently, as well as Machine Learning models.
  • Used Azure Data platform capabilities such as Azure Data Lake, Azure Data Factory, HDInsight, Azure SQL Server, Azure Machine Learning and Power BI to build huge Lambda systems.
  • Experience in working wif AWS Cloud services and SDKs to work wif services like AWS API Gateway, Lambda, S3, IAM, and EC2.
  • Using Linked Services/Datasets/Pipeline/, created Pipelines in ADF to extract, transform, and load data from various sources such as Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool, and backwards.
  • Experience in building, deploying, troubleshooting, and data extraction for huge number of records using Azure Data Factory (ADF).
  • Good Knowledge on Azure Synapse Analytics architecture and component integrations
  • Proactively monitored Synapse workload (Dedicated SQL pool and Spark pool) and engaged internal / external teams as appropriate to resolve errors / failures / throughput bottleneck
  • Configured and deployed EC2, VPC, S3, IAM, ELB, AWS Lambda, DynamoDB, CloudWatch, CloudFormation, AWS Auto Scaling, and Route53, among other AWS services.
  • Good knowledge on Core Java, J2EE, JDBC, Servlets, JSP, Exception Handling, Multithreading, EJB, XML, HTML5, CSS3, JavaScript, and AngularJS.
  • Experience in Google Cloud components, Google container builders and GCP client libraries and cloud SDK’s.
  • Worked wif a variety of databases, including Oracle, SQL Server, and MySQL, as well as written stored procedures, functions, joins, and triggers for various Data Models.
  • Experience in building and architecting multiple Data pipelines, end to end ETL and ELT process for Data ingestion and transformation in GCP.
  • Creating and monitoring batches and sessions using informatica Power center server.
  • Analyzed teh data sources and targets using Informatica data profiling option.
  • Design and develop ETL processes. Establish coding standards, perform code review and automate teh health of teh platform
  • Perform data analysis, develop database design, share, and get consensus on design, and tan develop teh stored procedures to fulfill teh database design
  • Manages end-user accounts and accessibility; provides technical expertise to end-users who create complex queries and reports
  • Experience wif teh Python libraries NumPy, Matplotlib, Pandas, Seaborn, and Plotly. Used PySpark, NumPy, and pandas on big datasets.
  • Extensive knowledge of relational and non-relational databases, including Oracle, PL/SQL, SQL Server, MySQL, and DB2.
  • Scripts and batch jobs were created to schedule various Hadoop programs. TensorFlow was used to train teh model from useful data.
  • Using Apache Kafka and Flume, we flooded HDFS wif massive amounts of data.
  • Installation and integration of Kafka wif Spark Streaming knowledge.
  • Experience wif Snowflake Multi-Cluster Warehouses, Snowflake Databases, Schemas, and Table Structures.
  • Able to handle numerous priorities and workloads, as well as quickly learn and adapt to new technologies and environments.
  • Experience in using Databricks in creation of notebooks and scheduling jobs.

TECHNICAL SKILLS

Hadoop Services: DFS, Map Reduce, Spark, YARN, Pig, Impala, Sqoop, Flume, Kafka, Stream-sets, NiFi, Hue

Hadoop Distribution: Cloudera, Hortonworks

Databases: HBase, Spark-Redis, Cassandra, Oracle, MySQL, Azure Cosmos, Mongo, Postgress, Teradata, Hive, Versant

Scheduling Tools: AutoSys, Oozie, Airflow

Cloud Computing Tools: AWS (Amazon EC2, Amazon EMR, LAMBDA, Amazon GLUE, Amazon S3, ATHENA), Azure (Azure Data Lake, Azure Data Factory, Azure Databricks, Azure SQL Database, Azure SQL data Warehouse), GCP.

Programming Languages: Python, Java, Scala, SQL, PL/SQL, Pig Latin, HiveQL, Unix, JavaScript, Shell Scripting.

Java & J2EE Technologies: Core Java, Servlets, Hibernate, Spring, Struts, JMS, EJB.

Operating Systems: UNIX, Windows, LINUX. Build Tools Jenkins, Maven, ANT.

Open-Source Software: Kubernetes, Docker, Git Hub, Rundeck

Development Tools: Eclipse, NetBeans, Microsoft SQL Studio, Toad.

PROFESSIONAL EXPERIENCE

Big Data Engineer

Confidential, Cincinnati, OH

Responsibilities:

  • Implemented Hadoop framework to capture user navigation across teh application to validate teh user interface and provide analytic feedback/result to teh UI team.
  • Loaded data into teh cluster from dynamically generated files using Flume and from relational database management systems using Sqoop.
  • Involved in performance tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and memory tuning.
  • Built a real time streaming pipeline utilizing Kafka, Spark Streaming and Redshift.
  • Developed logical and physical data flow models for Informatica ETL applications.
  • Worked on creation of customer Docker container images, tagging, and pushing of data images.
  • Written Hive queries for data analysis to meet teh business requirements.
  • Implemented and analyzed SQL query performance issues in databases.
  • Responsible for design development of Spark SQL Scripts based on Functional Requirements and Specifications.
  • Responsible for implementing various ML models both Supervised and Unsupervised models.
  • Hands on experience in loading data from UNIX file system to HDFS.
  • Experienced on loading and transforming of large sets of structured, semi structured and unstructured data from HBase through Sqoop and placed in HDFS for further processing.
  • Managing and scheduling Jobs on a Hadoop cluster using Oozie.
  • Involved in creating Hive tables, loading data, and running hive queries in those data.
  • Extensive Working knowledge of partitioned table, UDFs, performance tuning, compression-related properties, thrift server in Hive.
  • Worked on Microsoft Azure services like HDInsight clusters, BLOB, ADLS, Data Factory and Logic Apps and done POC on Azure Data Bricks.
  • Monitored Azure Synapse Analytics using Dynamic Management views to identify teh performance bottlenecks.
  • CreateCosmosscripts which pulls data from upstream structured streams and include business logic & transformations to meet teh requirements.
  • Worked wif Parquet file formats to process batch and interactive workloads.
  • Experience in developing spark applications using spark-SQL in databricks for data extraction, transformation, and aggregation from multiple file formats for analyzing and transforming teh data to uncover insights into teh customer usage patterns.
  • Extract Transform and Load data from Sources Systems to Azure Data Storage services using a combination of Azure Data Factory, T-SQL, Spark SQL, and U-SQL Azure Data Lake Analytics. Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing teh data in InAzure Databricks.
  • Created Pipelines in ADF using Linked Services/Datasets/Pipeline/ to Extract, Transform, and load data from different sources like Azure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards.
  • Developed JSON Scripts for deploying teh Pipeline in Azure Data Factory (ADF) dat process teh data using teh SQL Activity.
  • Used ELK cluster to organize error codes and log information for continuous monitoring spark applications and created alters in Kibana dashboard to trigger emails to production support team distribution list for any job failure.
  • Troubleshooting any build issue wif ELK and work towards teh solution.
  • Participated in problem resolving, change, release, and event management for ELK stack.
  • Designed end to end scalable architecture to solve business problems using various Azure Components like HDInsight, Data Factory, Data Lake, Storage and Machine Learning Studio. Developed JSON Scripts for deploying teh Pipeline in Azure Data Factory (ADF) dat process teh data using teh SQL Activity.
  • Worked alone on Google Cloud Migration project from Azure PaaS and converted their ETLs through Databricks and SSIS. Scheduled their jobs thru Databricks Jobs and SQL Server deployed packages wif SQL Server Agent.
  • Worked on Databricks in creation of multiple notebooks as per teh requirements and scheduling jobs using Apache Airflow.

Environment: Azure, ADF, Azure Synapse, Apache Hadoop, MapReduce, HDFS, CentOS 6.4, Spark, Kafka, HBase, Hive, Pig, Oozie, Flume, Java (jdk 1.6), Eclipse, ELK.

Data Engineer

Confidential, Rockledge, FL

Responsibilities:

  • Used Spark to import customer information data from Oracle database into HDFS for data processing along wif minor cleansing.
  • Developed MapReduce jobs to calculate teh total usage of data by commercial routers in different locations using Horton work distribution.
  • Involved in information gathering for new enhancements in Spark, Production support for field issues and label installs for Hive scripts and MapReduce jobs.
  • Managed GIT repositories for branching, merging, and tagging.
  • Used Maven to build rpms from source code in Scala checked out from GIT repository, wif Jenkins being teh Continuous Integration Server and Artifactory as repository manager.
  • Responsible for Setting up UNIX/Linux environments for various applications using shell scripting.
  • Set up scalability for application servers using command line interface for Setting up and administering DNS system in AWS using Route53.
  • Managed users and groups using teh Amazon Identity and Access Management (IAM).
  • In-depth understanding of teh principals and best practices ofSoftwareConfiguration
  • Designed and developed server-side applications on Linux platform in a fast-paced environment.
  • Translated customer business requirements into technical design documents, established specific solutions, and leading teh efforts including programming in Spark Scala and testing dat culminate in client acceptance of teh results.
  • Expertise in Object-Oriented Design (OOD) and end-to-endsoftwaredevelopment experience working on Scala coding and implementing mathematical models in Spark Analytics.
  • Softwaretranslating teh business requirements into Use Cases and Diagrams conducting reviews of Codes and test cases analyzing change requests enhancements managing release plans for business apps
  • Developed oozie workflows and scheduling jobs through Hue.
  • Used AWS Lambda to perform data validation, filtering, sorting and other transformations for every data change in a HBase table and load teh transformed data to RDS.
  • Loading data from different servers to S3 bucket and setting appropriate bucket permissions.
  • Configured routing to send JMS files to interact wif application for real time data using Kafka.
  • Managed Zookeeper for cluster co-ordination and Kafka Offset monitoring.
  • Optimized legacy queries to extract teh customer information from Oracle.
  • Reviewed HDFS usage and system design for future scalability and fault tolerance.
  • Managed and reviewed Hadoop log files using Spark to identify issues when job fails.

Environment: HDFS, Spark, Scala, Python, Shell Scripting, Hive, HBase, Oracle, MapReduce, Logstash, Jenkins, Versant, Java, Kafka, Horton works, GIT, ClearCase, Zookeeper, Ansible, AWS.

Big Data Engineer

Confidential

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop.
  • Used Spark for improving performance and optimization of teh existing algorithms in Hadoop using Spark Context, Spark Sessions, Spark-SQL, Data Frame, Pair RDD's, Spark YARN.
  • Designed and developed POCs in Spark using Scala to compare teh performance of Spark wif Map Reduce, Hive.
  • Involved in creating Hive tables and loading and analyzing data using hive queries.
  • Involved in design and development of an application using Hive (UDF).
  • Used HUE for Hive Query execution.
  • Used Hive QL to analyze teh partitioned and bucketed data, Executed Hive queries on Parquet tables stored in Hive to perform data analysis to meet teh business specification logic.
  • Developed spark jobs using SCALA and PYTHON on top of YARN/MR v2 for interactive and batch analysis.
  • Developed PYTHON script to start and end a job smoothly for a UC4 workflow.
  • Built real timedata pipelinesby developingKafkaproducers andSparkstreaming applications for consuming.
  • Written Hive queries on teh analyzed data for aggregation and reporting.
  • Implemented Spark Scripts using Scala, Spark SQL to access hive tables into spark for faster processing of data.
  • Developed end to end data processing pipelines dat begin wif receiving data using distributed messaging systemsKafkathrough persistence of data intoHBase.
  • Created a Python based web application using Python scripting for data processing, MySQL for teh database, and HTML/CSS/JQuery and HighCharts, matplotlib for data visualization of sales, tracking progress, identifying trends.
  • Developed various Python scripts to find vulnerabilities wif SQL Queries by doing SQL injection, permission checks and performance analysis.
  • Used Hive Context to integrate Hive metastore and Spark SQL for optimum performance.
  • Developed Sqoop Jobs to load data from RDBMS to external systems like HDFS and HIVE.
  • UsedInformaticaDesignertocreate,load,updatemappingsusingdifferenttransformationstomoveDataMartsinDataWarehouse.
  • PerformeddataanalysisanddataprofilingusingSQLandInformaticaDataExploreronsourcessystemsincludingOracleandTeradata.
  • AnalyzedtheETLprocesscreatedinInformatica,developedSQLqueriesandtableswhichreplic-atedtheETL,andcomparedtheSQLresultswiftheInformaticatables.
  • WroteSQLscriptstotestmappingsanddevelopedTraceabilityMatrixtotestscriptsforanyChangeControlinrequirementsleadingtotestcaseupdate.
  • ReviewedbasicSQLqueriesandeditedinner,left,andrightjoinsinTableauDesktopbyconnectinglive/dynamicandstaticdatasets.
  • CreatedvisuallyimpactfuldashboardsinExcelandTableaufordatareportingwifrespecttocu-stomer,policyinformationandanalyzedtoidentifythekeymetrics.
  • Launching Amazon EC2 Cloud Instances using Amazon Images (Linux/ Ubuntu) and configuring launched instances wif respect to specific applications.
  • Added support for AWS S3 and RDS to host static /media files and teh database into amazon cloud.
  • Ingested data from disparate data sources using a combination of SQL, Google Analytics API, and salesforce API using python to create data views to be used in BI tools like Tableau.
  • Worked in Agile Iterative sessions to create Hadoop Data Lake for teh client
  • Developed Stored Procedures and Functions, Views for teh Oracle database PL/SQL.
  • Responsible for generating actionable insights from complex data to drive real business results for various application teams and worked in Agile Methodology projects extensively.

Environment: Amazon services, Spark SQL, PYTHON, CDH, HDFS, Hive, Pig, Apache Sqoop, Java (JDK SSE 6, 7), Scala, Shell scripting, Linux, MYSQL Oracle Enterprise DB, Eclipse, Oracle, Git, Oozie, Informatica, Tableau, MySQL, Soap, and Agile Methodologies.

Java Developer

Confidential

Responsibilities:

  • Responsible for acquiring, assessing, and translating requirements into technical specifications.
  • Created sequence and class diagrams.
  • Created presentation layer wif JSP, Java, HTML, and JavaScript.
  • Used Spring Core Annotations for Dependency Injection.
  • To enable dynamic fetching and displaying of diverse table data using JSF tag libraries, I designed and created a ‘Convention Based Coding’ using Hibernates persistence framework and O-R mapping capability.
  • For database connectivity and accessing teh session for database transactions, I designed and built Hibernate configuration and teh session-per-request design pattern. For retrieving and saving data in databases, I used HQL and SQL.
  • Contributed to teh design and implementation of database schema and Entity-Relationship diagrams for teh application's backend Oracle database tables.
  • Used Apache Axis to implement web services.
  • Designed and developed Oracle Stored Procedures and Triggers to meet teh application's requirements.
  • Created sophisticated SQL queries to pull data from teh database SOAP web service interfaces implemented in Java were designed and built.
  • For teh build process, Apache Ant was used.

Environment: Java, JDK 1.5, Servlets, Hibernate, Ajax, Oracle 10g, Eclipse, Apache Ant, Web Services (SOAP), Apache Axis, Apache Ant, Web Logic Server, JavaScript, HTML, CSS, XML.

We'd love your feedback!