We provide IT Staff Augmentation Services!

Senior Big Data Engineer Resume

4.00/5 (Submit Your Rating)

Weehawken, NJ

SUMMARY

  • Overall 8+ years of professional experience in Information Technology and expertise in BIGDATA using HADOOP framework and Analysis, Design, Development, Testing, Documentation, Deployment and Integration using SQL and Big Data technologies.
  • Experienced in managing Hadoop clusters and services using Cloudera Manager.
  • Expertise in using various Hadoop infrastructures such as Map Reduce, Pig, Hive, Zookeeper, Hbase, Sqoop, Oozie, Airflow, Snowflake, Flume, Drill and spark for data storage and analysis.
  • Experience in Developing Spark applications using Spark - SQL in Databricks for data extraction, transformation and aggregation from multiple file formats for analyzing & transforming the data to uncover insights into the customer usage patterns.
  • Implemented various algorithms for analytics using Cassandra with Spark and Scala.
  • Experienced in Creating Vizboards for data visualization in Platfora for real - time dashboard on Hadoop.
  • Collected logs data from various sources and integrated in to HDFS using Flume.
  • Designed and implemented a product search service using Apache Solr.
  • Experience in developing custom UDFs for Pig and Hive to in corporate methods and functionality of Python/Java into Pig Latin and HQL(HiveQL) and Used UDFs from Piggybank UDF Repository.
  • Experienced in running query - usingImpalaand used BI tools to run ad-hoc queries directly on Hadoop.
  • Good experience in Oozie Framework and Automating daily import jobs.
  • Expertise with Big data on AWS cloud services i.e. EC2, S3, Auto Scaling, Glue, Lambda, Cloud Watch, Cloud Formation, Anthena, DynamoDB and RedShift
  • Strong experience in core Java,Scala, SQL, PL/SQL and Restful web services.
  • Extensive knowledge in various reporting objects like Facts, Attributes, Hierarchies, Transformations, filters, prompts, calculated fields, Sets, Groups, Parameters etc., in Tableau experience in working with Flume and NiFi for loading log files into Hadoop.
  • Strong experience in Snowflake.
  • Experienced in troubleshooting errors in Hbase Shell/API, Pig, Hive and map Reduce.
  • Highly experienced in importing and exporting data between HDFS and Relational Database Management systems using Sqoop.
  • Good knowledge in querying data from Cassandraf or searching grouping and sorting.
  • Extensive hands-on experience in using distributed computing architectures such as AWS products (e.g. EC2, Redshift, EMR, and Elastic search), Hadoop, Python, Spark and effective use of Azure SQL Database, MapReduce, Hive, SQL and PySpark to solve big data type problems.
  • Ability to work effectively in cross-functional team environments, excellent communication, and interpersonal skills. Involved in converting Hive/SQL queries into Spark transformations using Spark Data frames and Scala.
  • Developed custom Kafka producer and consumer for different publishing and subscribing to Kafka topics.
  • Good working experience on Spark (spark streaming, spark SQL) with Scala and Kafka. Worked on reading multiple data formats on HDFS using Scala.
  • Good understanding of NoSQL Data bases and hands on work experience in writing applications on No SQL data bases like Cassandra and Mongo DB.
  • Creative skills in developing elegant solutions to challenges related to pipeline engineering
  • Good Understanding of Azure Big data technologies like Azure Data Lake Analytics, Azure Data Lake Store, Azure Data Factory, Azure Databricks, and created POC in moving the data from flat files and SQL Server using U-SQL jobs.
  • Experience in using Kafka and Kafka brokers to initiate spark context and processing livestreaming.
  • Worked on various programming languages using IDEs like Eclipse, NetBeans, and Intellij, Putty, GIT.

TECHNICAL SKILLS

Big Data Ecosystem: Hadoop, MapReduce, Pig, Hive, YARN, Kafka, Flume, Sqoop, Impala, Oozie, Zookeeper, Spark, Ambari, Snowflake, Airflow, Mahout, MongoDB, Cassandra, Avro, Storm, Parquet and Snappy.

Hadoop Distributions: Cloudera (CDH3, CDH4, and CDH5), Hortonworks, MapR and Apache

Languages: Java, Python, Jruby, SQL, HTML, DHTML, Scala, JavaScript, XML and C/C++

No SQL Databases: Cassandra, MongoDB and HBase

Java Technologies: Servlets, JavaBeans, JSP, JDBC, JNDI, EJB and struts

XML Technologies: XML, XSD, DTD, JAXP (SAX, DOM), JAXB

Development Methodology: Agile, waterfall

Development / Build Tools: Eclipse, Ant, Maven, IntelliJ, JUNIT and log4J

Frameworks: Struts, spring and Hibernate

App/Web servers: WebSphere, WebLogic, JBoss and Tomcat

SQL Databases: MySQL, MS SQL, PL/SQL, and Oracle

Cloud Technologies: AWS, Azure

PROFESSIONAL EXPERIENCE

Confidential, Weehawken, NJ

Senior Big Data Engineer

Responsibilities:

  • Expertise AWS Lambda function and API Gateway, to submit data via API Gateway that is accessible via Lambda function.
  • Managed configuration of Web App and Deploy to AWS cloud server through Chef.
  • Making a data pipelining with help Data Fabric job, SQOOP, SPARK, Scala and KAFKA. Parallel working in data side oracle and MYSQL server for data designing to source to target.
  • Developed PIG Latin scripts for the analysis of semi structured data.
  • Involved in designing and deployment of Hadoop cluster and different Big Data analytic tools including Pig, Hive, HBase, Oozie, Zookeeper, SQOOP, flume, Spark, Impala, and Cassandra with Horton work Distribution.
  • Developed AWS Cloud Formation templates and set up Auto scaling for EC2 instances.
  • Involved in Importing and exporting data from HDFS using Sqoop, resolution of access issues, performance issues and Patch/upgrade related issues.
  • Responsible for estimating the cluster size, monitoring and troubleshooting of the Spark databricks cluster.
  • Installed Airflow and created a database in PostgreSQL to store metadata from Airflow.
  • Configured documents which allow Airflow to communicate to its PostgreSQL database.
  • Experience in working with Map Reduce programs using Apache Hadoop for working with Big Data
  • Involved in creating HiveQL on HBase tables and importing efficient work order data into Hive tables
  • Extending Hive and Pig core functionality by writing customUDFs, UDTF and UDAFs.
  • Used Oozie and Zookeeper operational services for coordinating cluster and scheduling workflows.
  • Write programs using Spark to move data from Storage input location to output location by running data loading, validation, and transformation to the data
  • Created data sharing between two snowflake accounts.
  • Developed Kafka consumer API in Scala for consuming data from Kafka topics.
  • Created data pipeline for different events of ingestion, aggregation and load consumer response data in AWS S3 bucket into Hive external tables in HDFS location to serve as feed for tableau dashboards
  • Developed Java Map Reduce programs for the analysis of sample log file stored in cluster.
  • Working experience with data streaming process with Kafka, Apache Spark, Hive.
  • Responsible for Design, Development, and testing of the database and Developed Stored Procedures, Views, and Triggers
  • Working as Developer in hive and impala for more parallel processing data in Cloudera systems.
  • Working in big data technologies like spark, Scala, Hive, Hadoop cluster (Cloudera platform).
  • Design & implement Spark SQL tables, Hive scripts job with stone branch for scheduling and create work flow and task flow.
  • Excellent understanding of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm.
  • Imported documents into HDFS, HBase and creating HAR files.
  • Expertise in data transformation & analysis usingSPARK,PIG, HIVE
  • We generally used partitions and bucketing for data in hive to get query faster. This part of hive optimization
  • Converted Talend Joblets to support the snowflake functionality.
  • Responsible to manage data coming from different sources through Kafka.
  • Good Exposure on Map Reduce programming using Java, PIG Latin Scripting and Distributed Application and HDFS.
  • Created instances in AWS as well as worked on migration to AWS from data center.
  • Developed Python-based API (RESTful Web Service) to track revenue and perform revenue analysis.
  • Compiling and validating data from all departments and Presenting to Director Operation.
  • Working on designing the Map Reduce and Yarn flow and writing Map Reduce scripts, performance tuning and debugging.
  • Utilized spark SQL to load data from AWS S3 to snowflake tables using data bricks.
  • Working in relational SQL and NoSQL databases, including Oracle, Hive, Sqoop and HBase
  • Experienced on implementation of a log producer in Scala that watches for application logs, transform incremental log and sends them to a Kafka and Zookeeper based log collection platform.
  • Created Notebooks using Databricks, Scala and spark and capturing the data from Delta tables in Delta lakes.
  • Migrated Map reduce jobs to Spark jobs to achieve better performance.
  • Handled importing of data from various data sources, performed transformations using Hive, MapReduce, loaded data into HDFS and extracted data from MYSQL into HDFS vice-versa using Sqoop.
  • Used Scala function, dictionary and data structure (array, list, map) for better code reusability
  • Based on Development, we need to do the Unit Testing.
  • Migrated an existing on-premises application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage, experienced in Maintaining the Hadoop cluster on AWS EMR
  • Utilized Spark SQL API in PySpark to extract and load data and perform SQL queries.
  • Worked on developing Pyspark script to encrypting the raw data by using Hashing algorithms concepts on client specified columns.
  • Unit tested the data between Redshift and Snowflake.
  • Responsible for distributed applications across hybrid AWS
  • Implemented data ingestion and handling clusters in real time processing usingKafka.
  • Creating datamodel that correlates all the metrics and gives a valuable output.
  • Worked on the tuning of SQL Queries to bring down run time by working on Indexes and Execution Plan.
  • Performing ETL testing activities like running the Jobs, Extracting the data using necessary queries from database transform, and upload into the Data warehouse servers.
  • Pre-processing using Hive and Pig.

Environment: HDFS, Hive, Spark, Airflow, AWS, EC2, S3, Lambda, Auto Scaling, Cloud Watch, Linux, Kafka, python, Stone branch, Cloudera, Databricks, Talend, Snowflake, Oracle12c, PL/SQL, Unix, Json and Parquet File systems.

Confidential, New York, NY

Big Data Engineer/Hadoop Developer

Responsibilities:

  • Involved in complete Big Data flow of the application starting from data ingestion upstream to HDFS, processing the data in HDFS and analyzing the data and involved.
  • Created Partitioned Hive tables and worked on them using HiveQL.
  • Developed Json Scripts for deploying the Pipeline in Azure Data Factory (ADF) that process the data using the Cosmos Activity.
  • Responsible for data services and data movement infrastructures
  • Experienced in ETL concepts, building ETL solutions and Data modeling
  • Used OOZIE Operational Services for batch processing and scheduling workflows dynamically
  • Worked with SCRUM team in delivering agreed user stories on time for every Sprint.
  • Experience in developing customized UDF’s in Python to extend Hive and Pig Latin functionality.
  • Created Pipelines inADFusingLinked Services/Datasets/Pipeline/ to Extract, Transform and load data from different sources likeAzure SQL, Blob storage, Azure SQL Data warehouse, write-back tool and backwards.
  • Redesigned the Views in snowflake to increase the performance.
  • Configured Flume to extract the data from the web server output files to load into HDFS.
  • Used Flume to collect, aggregate and store the web log data from different sources like web servers, mobile and network devices and pushed into HDFS.
  • Develop batch processing solutions by using Data Factory and Azure Data bricks
  • Implement Azure Data bricks clusters, notebooks, jobs and auto scaling.
  • Develop solutions to leverage ETL tools and identify opportunities for process improvements using Informatica and Python
  • Worked on architecting the ETL transformation layers and writing spark jobs to do the processing.
  • Experience in developing MapReduce Programs using Apache Hadoop for analyzing the big data as per the requirement.
  • Implemented Kafka Security Features using SSL and without Kerberos. Further with more grain-fines Security I set up Kerberos to have users and groups this will enable more advanced security features
  • Measured Efficiency of Hadoop/Hive environment ensuring SLA is met
  • Analyzed the system for new enhancements/functionalities and perform Impact analysis of the application for implementing ETL changes
  • Designed and implemented by configuring Topics in new Kafka cluster in all environment
  • Developing Json Scripts for deploying the Pipeline in Azure Data Factory (ADF) that process the data using the Cosmos Activity.
  • Developed Airflow DAGs in python by importing the Airflow libraries.
  • Used windows Azure SQL reporting services to create reports with tables, charts and maps.
  • Wrote Flume configuration files for importing streaming log data into HBase with Flume.
  • Imported several transactional logs from web servers with Flume to ingest the data into HDFS.
  • Using Flume and Spool directory for loading the data from local system (LFS) to HDFS.
  • Involved in designing the row key in HBase to store Text and JSON as key values in HBase table and designed row key in such a way to get/scan it in a sorted order.
  • Creating Reports in Looker based on Snowflake Connections
  • Demonstrated expert level technical capabilities in areas of Azure Batch and Interactive solutions, Azure Machine learning solutions and operationalizing end to end Azure Cloud Analytics solutions.
  • Designed end to end scalable architecture to solve business problems using various Azure Components like HDInsight, Data Factory, Data Lake, Storage and Machine Learning Studio.
  • Worked on migrating MapReduce programs into Spark transformations using Spark and Scala.
  • Designed changes to transform current Hadoop jobs to HBase.
  • Partnered with ETL developers to ensure that data is well cleaned and the data warehouse is up-to-date for reporting purpose by Pig.

Environment: Hadoop YARN, Spark, Spark Streaming, Airflow, MapReduce, Spark SQL,Kafka, Scala, Azure, Data Bricks, Data Factory, Data Lake, Python, Hive, Sqoop, Impala, Tableau, Talend, Snowflake, Oozie, Control-M, HBase, Java, Oracle 12c, Linux

Confidential, Chicago, IL

Data Engineer/ Hadoop Developer

Responsibilities:

  • Worked on developing ETL processes (Data Stage Open Studio) to load data from multiple data sources to HDFS using FLUME and SQOOP, and performed structural modifications using Map Reduce, HIVE.
  • Worked with HIVE data warehouse infrastructure-creating tables, data distribution by implementing partitioning and bucketing, writing and optimizing the HQL queries.
  • Built and implemented automated procedures to split large files into smaller batches of data to facilitate FTP transfer which reduced 60% of execution time.
  • Optimized and tuned ETL processes & SQL Queries for better performance.
  • Developed PIG UDFs for manipulating the data according to Business Requirements and also worked on developing custom PIG Loaders.
  • Developed Python scripts to take backup of EBS volumes using AWS Lambda and Cloud Watch.
  • Transformed the data using AWS Glue dynamic frames with PySpark; cataloged the transformed the data using Crawlers and scheduled the job and crawler using workflow feature
  • Worked on installing cluster, commissioning & decommissioning of data node, name node recovery, capacity planning, and slots configuration.
  • Developed and deployed stacks using AWS Cloud Formation Templates (CFT) and AWS Terraform.
  • Developed data pipeline programs with Spark Scala APIs, data aggregations with Hive, and formatting data (JSON) for visualization, and generating.
  • Developing Spark scripts, UDFS using both Spark DSL and Spark SQL query for data aggregation, querying, and writing data back into RDBMS through Sqoop.
  • Interacted with business partners, Business Analysts and product owner to understand requirements and build scalable distributed data solutions using Hadoop ecosystem.
  • Developed Spark Streaming programs to process near real time data from Kafka, and process data with both stateless and state full transformations.
  • Involved in converting Hive/SQL queries into transformations using Python
  • Experience in report writing using SQL Server Reporting Services (SSRS) and creating various types of reports like drill down, Parameterized, Cascading, Conditional, Table, Matrix, Chart and Sub Reports.
  • Used DataStax Spark connector which is used to store the data into Cassandra database or get the data from Cassandra database.
  • Wrote Oozie scripts and setting up workflow using Apache Oozie workflow engine for managing and scheduling Hadoop jobs.
  • Worked on implementation of a log producer in Scala that watches for application logs, transform incremental log and sends them to a Kafka and Zookeeper based log collection platform.
  • Used Hive to analyze data ingested into HBase by using Hive-HBase integration and compute various metrics for reporting on the dashboard.
  • Written multiple MapReduce Jobs using Java API, Pig and Hive for data extraction, transformation and aggregation from multiple file formats including Parquet, Avro, XML, JSON, CSV, ORCFILE and other compressed file formats Codecs like gZip, Snappy, Lzo.
  • Strong understanding of Partitioning, bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance.

Environment: Apache Spark, Map Reduce, Apache Pig, Python, SQL, Java, SSRS, HBase, AWS, Cassandra, PySpark, Apache Kafka, HIVE, SQOOP, FLUME, Apache Oozie, Zookeeper, ETL, UDF.

Confidential

Data Engineer

Responsibilities:

  • Performed Data Cleaning, features scaling, features engineering using pandas and NumPy packages in python and build models using deep learning frameworks
  • Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
  • Researched and downloaded jars for Spark-avro programming.
  • Developed a PySpark program that writes dataframes to HDFS as avro files.
  • Utilized Spark's parallel processing capabilities to ingest data.
  • Transformed batch data from several tables containing tens of thousands of records from SQL Server, MySQL, PostgreSQL, and csv file datasets into data frames using PySpark.
  • Utilized Airflow to schedule automatically trigger and execute data ingestion pipeline.
  • Implemented clustering techniques like DBSCAN, K-means, K-means++ and Hierarchical clustering for customer profiling to design insurance plans according to their behavior pattern.
  • Worked with Customer Churn Models including Random forest regression, lasso regression along with pre-processing of the data.
  • Created and executed HQL scripts that creates external tables in a raw layer database in Hive.
  • Developed a Script that copies avro formatted data from HDFS to External tables in raw layer.
  • Created PySpark code that uses Spark SQL to generate dataframes from avro formatted raw layer and writes them to data service layer internal tables as orc format.

Environment: Spark, Redshift, Python, HDFS, Hive, Pig, Sqoop, Scala, Kafka, Shell scripting, Linux, Jenkins, Eclipse, Git, Oozie, Spark, Cloudera, Oracle 10g, PL/SQL, Unix.

Confidential

SQL Developer

Responsibilities:

  • Involved in complete Software Development Lifecycle (SDLC).
  • Worked on different dataflow and control flow task, for loop container, sequence container, script task, executes SQL task and Package configuration.
  • Extensive use of Expressions, Variables, Row Count in SSIS packages
  • Created SSIS packages to pull data from SQL Server and exported to Excel Spreadsheets and vice versa.
  • Created batch jobs and configuration files to create automated process using SSIS.
  • Data validation and cleansing of staged input records was performed before loading into Data Warehouse
  • Automated the process of extracting the various files like flat/excel files from various sources like FTP and SFTP (Secure FTP).
  • Deploying and scheduling reports using SSRS to generate daily, weekly, monthly and quarterly reports.
  • Loading data from various sources like OLEDB, flat files to SQL Server database Using SSIS Packages and created data mappings to load the data from source to destination.
  • Built SSIS packages, to fetch file from remote location like FTP and SFTP, decrypt it, transform it, mart it to data warehouse and provide proper error handling and alerting

Environment: MS SQL Server, SQL Server Business Intelligence Development Studio, SSIS, SSRS, Report Builder, Office, Excel, Flat Files, T-SQL.

We'd love your feedback!