We provide IT Staff Augmentation Services!

Big Data Developer Resume

0/5 (Submit Your Rating)

Austin, TX

SUMMARY

  • Around 7 years of experience as a Data Engineer and extensively worked with designing, developing, and implementing Data models for enterprise - level applications and BI solutions.
  • Experience in designing and building Data Management Lifecycle covering Data Ingestion, Data integration, Data consumption, Data delivery, and integration Reporting, Analytics, and System-System integration.
  • Proficient in Big Data environment and Hands-on experience in utilizing Hadoop environment components for large-scale data processing including structured and semi-structured data.
  • Strong experience with all phases including Requirement Analysis, Design, Coding, Testing, Support, and Documentation.
  • Extensive experience with Azure cloud technologies like Azure Data Lake Storage, Azure Data Factory, Azure SQL, Azure Data Warehouse, Azure Synapse Analytical, Azure Analytical Services, Azure HDInsight, and Databricks.
  • Experience with monitoring the web services using Hadoop and Spark for controlling the applications and analysing their operation and performance.
  • Experienced in Python data manipulation for loading and extraction as well as with Python libraries such as NumPy, Pandas, and SciPy for data analysis and numerical computations.
  • Good knowledge and experience with NoSQL databases like HBase, Cassandra, and MongoDB and SQL databases like Teradata, Oracle, PostgreSQL, and SQL Server.
  • Experience in the development and design of various scalable systems using Hadoop technologies in various environments and analysing data using MapReduce, Hive, and PIG.
  • Hands-on use of Spark and Scala to compare the performance of Spark with Hive and SQL, and Spark SQL to manipulate Data Frames in Scala.
  • Strong knowledge in working with ETL methods for data extraction, transformation, and loading in corporate- wide ETL Solutions and Data Warehouse tools for reporting and data analysis.
  • Hands-on experience in designing and implementing data engineering pipelines and analysing data using Hadoop ecosystem tools like HDFS, Spark, Sqoop, Hive, Flume, Kafka, Impala, PySpark, Oozie, and HBase.
  • Strong experience with all phases including Requirement Analysis, Design, Coding, Testing, Support, and Documentation.
  • Experience with different ETL tool environments like SSIS, Informatica, and reporting tool environments like SQL Server Reporting Services, and Business Objects.
  • Experience in deployment of applications and scripting using the Unix/Linux Shell scripting.
  • Solid knowledge of Data Marts, Operational Data Store, OLAP, Dimensional Data Modelling with Star Schema Modelling, Snowflake Modelling for Dimensions Tables using Analysis Services.
  • Extensive experience with various databases like Teradata, MongoDB, Cassandra DB, MySQL, Oracle, and SQL Server.
  • Participated in the design and development of data solutions that use Azure Data Factory and Azure Synapse Analytics to support the Business Intelligence team to deliver actionable insights.
  • Worked with the Data Science team to construct Gold tables based on business requirements for premium members providing further athlete performance and analysis and created managed Hive tables in Databricks.
  • Experience in Creating Teradata SQL scripts using OLAP functions like rank and rank over to improve the query performance while pulling the data from large tables.
  • Strong Experience in working with Databases like Teradata and proficiency in writing complex SQL, PL/SQL for creating tables, views, indexes, stored procedures, and functions.
  • Knowledge and experience with Continuous Integration and Continuous Deployment using containerization technologies like Docker and Jenkins.
  • Excellent working experience in Agile/Scrum development and Waterfall project execution methodologies

TECHNICAL SKILLS

Hadoop/Big Data ecosystems: HDFS, Map Reduce, Sqoop, Flume, Pig, Hive, Oozie, Impala, Zookeeper and Cloudera Manager, Zookeeper, Spark, Scala

NoSQL Database: HBase, Cassandra

Tools: and IDE Eclipse, NetBeans, Toad, Putty, Maven, DB Visualizer, VS Code, Qlik Sense, Qlik View

Languages: SQL, PL/SQL, JAVA, Scala, Python

Databases: Oracle, SQL Server, MySQL, DB2, PostgreSQL, Teradata

Tracking Tools and Control: SVN, GIT, Maven

ETL Tools: OFSAA, IBM DataStage

Cloud Technologies: Azure, AWS

PROFESSIONAL EXPERIENCE:

Confidential, Austin, Tx

Big Data Developer

Responsibilities:

  • Research and recommend a suitable technology stack for Hadoop migration considering current enterprise architecture.
  • Expertise in writingHadoopJobsfor analyzing structured and unstructured data usingHDFS,Hive, HBase, Pig, Spark, Kafka, Scala, Oozie, andTalend ETL.
  • Extensively usedSparkstack to develop preprocessing job which includes RDD, Datasets, and Data frame APIs to transform the data for upstream consumption.
  • Build and maintain scalable data pipelines using theHadoop ecosystemand other open-source components likeHive, andHBase.
  • Good understanding ofNoSQL Databases and hands-on work experience in writing applications on No SQL databases likeCassandraandMongo DB.
  • Good knowledge in querying data fromCassandrafor searching grouping and sorting.
  • Involved in various NoSQL databases likeHBase, Cassandrain implementing and integration.
  • Involved in developing of JDBCDAOs and DTOs, access of advanced SQL and PL/SQL stored procedures on database systems using spring templates.
  • Onboarding of new data into Splunk. Troubleshooting Splunk and optimizing performance.
  • Developed multiple MapReduce jobs in java for data cleaning and pre-processing.
  • Worked on extracting and enriching relational databases data between multiple tables using joins inSpark.
  • Worked on writing APIs to load the processed data to Hive tables.
  • ExperiencedScalain using andspark streamingandAkkafor ongoing transactions for customers.
  • Replaced the existing Map Reduce programs intoSparkapplication using Scala.
  • Developed the Hive UDF's to handle data quality and create filtered datasets for further processing
  • Experienced in writing Sqoop scripts to import data into Hive/HDFS from RDBMS.
  • Applied Azure Data Factory to ingest raw data from several data sources and store it in Azure Data Lake Storage Gen2 for downstream processing.
  • Created dynamic pipeline with stored procedures or data flow in Azure Data Factory to update different dimensional tables for reducing the number of activities and easier maintenance.
  • Used Azure Databricks (PySpark) to exact raw data and transform it into bronze by adding metadata, then cleansed, validated, and enriched bronze data to silver/gold tables and persisted them in ADLS gen 2 based Delta Lake
  • Ingested real-time streaming sports game pipelines with Event Hubs and processed the athlete score and performance in real-time by Azure Stream Analytics for providing simple data to software engineers, which is displayed on the mobile application.
  • Conducted daily batch incremental load by Azure Data Factory to update the data for customer use. · Implemented Slowly Changing Dimension Type 2 in Delta Lake as the update strategy to keep the full historical player and team profiles.
  • Worked with the Data Science team to construct Gold tables based on business requirements for premium members providing further athlete performance and analysis and created managed Hive tables in Databricks
  • SupportedMap Reduce Programsthat are running on the cluster and wroteMapReduce jobsusingJava API.
  • Optimized Hive QL scripts by using execution engines like Tez, Spark.
  • Developed Hive queries to analyze the data in HDFS to identify issues and behavioral patterns.
  • Able to use Python Pandas, NumPy modules for Data analysis, Data scraping, and parsing.
  • Deployed applications using Jenkins’s framework integrating Git- version control with it.
  • Participated in production support regularly to support the Analytics platform
  • Participated in the design and development of data solutions that use Azure Data Factory and Azure Synapse Analytics to support the Business Intelligence team to deliver actionable insights.
  • Used Git and Azure DevOps to perform version control, peer review, and enhance collaboration.
  • Worked under Agile Scrum methodology with 2-week sprints, and participated in stand-up meetings, sprint planning, retrospective meetings, and code reviews

Environment: Hadoop, HDFS, AWS, Hive, BigQuery, Spark SQL, GCP, MapReduce,Sparkstreaming, Abinitio, Sqoop, Oozie, Jupiter Notebook, Docker, Kafka,Spark, Scala, Talend, Shell Scripting.

Confidential

Big Data Developer

Responsibilities:

  • Performed advanced procedures like text analytics and processing using the in-memory computing capabilities ofSpark.
  • DevelopedSparkcode using Scala andSpark-SQL for faster processing and testing.
  • Worked onSparkSQL for joining multi hive tables and write them to a final hive table.
  • Configured Spark Streaming to receive real-time data from the Apache Kafka and store the stream data toHDFSusingScala.
  • Worked on scalable distributed data system using Hadoop ecosystem in AWS EMR and MapR (MapR data platform).
  • Developed prototypeSparkapplications usingSpark-Core, Spark SQL, Data Frame APIand developed several custom User-defined functions inHive & Pig using Java & python
  • Importing the data intoSparkfromKafkaConsumer group usingSpark Streaming APIs.
  • Good Knowledge of reporting and data visualization tools like Oracle Data Visualization Desktop, Tableau, and Grafana.
  • Worked on Google Cloud Platform (GCP) Services like Vision API, Instances.
  • Used Informatica Cloud Data Integration for global, Data Cloud Architect, Azure, Ansible, Jenkins, Docker, Kubernetes, DevOps, Automation, CI/CD, distributed data warehouse, and analytics projects.
  • Utilized Azure Data Factory to create, Data Cloud Architect, Azure, Ansible, Jenkins, Docker, Kubernetes, DevOps, Automation, CI/CD, schedule and manage data pipelines.
  • Developed a POC for project migration from on prem Hadoop MapR system to GCP/Snowflake.
  • Build cluster on AWS environment using EMR using S3, EC2, and Redshift.
  • Built dashboards and visualizations on top of MapR-DB and Hive using Oracle data visualizer desktop. Built real-time visualizations on top of Open TSDB using Grafana.
  • Used DataStax Cassandra along with Pentaho for reporting.
  • Designed, configured, and deployed Amazon Web Services (AWS) for a multitude of applications utilizing the AWS stack (Including EC2, Glue, Data pipeline EMR, SNS, S3, RDS, Cloud Watch, SQS, IAM), focusing on high-availability, fault tolerance, and auto-scaling.
  • Worked with AWS Glue jobs to transform data to a format that optimizes query performance for Athena.
  • Working on new system architecture to replace client's current crediteTradingplatform usingFlink.
  • ImplementedSparkRDD transformations to Map business analysis and apply actions on top of transformations.
  • CreatedSparkjobs to do lighting speed analytics over theSparkcluster.
  • EvaluatedSpark's performance vs Impala on transactional data.
  • Experienced in developing scripts for doing transformations usingScala.
  • Experienced in creating data pipeline integratingKafkawithspark streamingapplication usedScalafor writing applications.
  • Developed customizedUDFs in Javafor extending Pig and Hive functionality.
  • Usedspark SQLfor reading data from external sources and processes the data usingScalacomputation framework.
  • UsedSparktransformations and aggregations to perform min, max, and average on transactional data.
  • Extracted files from databases through Sqoop and placed in HDFS and processed through spark.
  • Experienced in migrating Hive QL into Impala to minimize query response time.
  • Experience using Impala for data processing on top of HIVE for better utilization.
  • Developed and optimizedPigandHiveUDFsto implement the methods and functionality of Javaas required.
  • Wrote queries UsingCassandra CQLto create, alter, insert and delete elements.
  • Hands-on experience working onNoSQLdatabases includingHBase,MongoDB,Cassandra, and its integration withHadoop cluster.
  • ImplementedAWS EC2, Key Pairs, Security Groups, Auto Scaling, ELB, SQS, and SNSusingAWS APIand exposed as theRestful Web services andimplementedReporting, Notification servicesusingAWS API.
  • Performed querying of both managed and external tables created by Hive using Impala.
  • Developed Impala scripts for end-user/analyst requirements for Adhoc analysis.
  • Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
  • UsedSparkAPI over Cloudera Hadoop YARN to perform analytics on data in Hive.
  • UsedHTML, CSS, XML, JavaScriptandJSPfor interactive cross browser functionality and complex user interface.
  • Experience in writingcustom UDFsfor Hive to in corporate methods and functionality of Java into andHQLHIVESQL.
  • Collected data usingSparkStreaming from AWSS3 bucket in near-real-time and performs necessary Transformations and Aggregations to build the data model and persists the data in HDFS.
  • Responsible for creating Hive tables, loading with data, and writing Hive queries.
  • Optimized Hive QL by using execution engines like Tez,Spark.
  • Responsible for creating mappings and workflows to extract and load data from relational databases, flat file sources, and legacy systems using Abinitio.
  • Fetch and generate monthly reports, Visualization those reports using Tableau.
  • Used Oozie Workflow engine to run multiple Hive jobs.

Environment: Hadoop, Cloudera, Flume, HBase, GCP, HDFS, MapReduce, YARN, Hive, Sqoop, Oozie, Tableau, Abinitio, JUnit, agile methodologies, UNIX

Confidential

Hadoop/ETL Developer

Responsibilities:

  • Implemented ETL Abinitio designs and processes for a load of data from the sources to the target warehouse.
  • Experience in developing scalable real-time applications for ingesting clickstream data using Kafka Streams and Spark Streaming.
  • Configured several nodes EC2 Hadoop cluster to transfer the data from S3 to HDFS and vice-versa and to direct input and output to the Hadoop MapReduce framework.
  • Experience with different table structure, file formats, partitioning and bucketing concepts in Hive.
  • Experienced in converting Hive scripts into spark using Scala and optimized spark jobs.
  • Developed ETL Processes in AWS Glue to migrate Campaign data from external sources like S3 ORC/Parquet/Text Files into AWS Redshift.
  • Experience in developing Spark applications using Spark-SQL in Databricks for data extraction, transformation, and aggregation from multiple file formats for Analyzing & transforming the data to uncover insights into the customer usage patterns.
  • Create PySpark frame to bring data from DB2 to Amazon S3.
  • Pushed application logs and data streams logs to Kibana server for monitoring and alerting purpose.
  • Developed optimized and tuned ETL operations in Hive and Spark scripts using techniques such as partitioning, bucketing, vectorization, serialization, configuring memory and number of executors.
  • Implemented cloud integrations to GCP and Azure for bi-directional flow setups for data migrations.
  • Experience in developing Spark applications using Spark-SQL in Databricks for data extraction, transformation, and aggregation from multiple file formats for Analyzing & transforming the data to uncover insights into the customer usage patterns.
  • Pushed application logs and data streams logs to Kibana server for monitoring and alerting purpose.
  • Experience designing solutions in Azure tools like Azure Data Factory, Azure Data Lake, SQL DWH, Azure SQL & Azure SQL Data Warehouse, Azure Functions.
  • Worked on migrating data from HDFS to Azure HD Insights and Azure Databricks.
  • Migrated existing processes and data from our on-premises SQL Server and other environments to Azure Data Lake.
  • Developed and Tuned Spark Streaming application using Scala for processing data from Kafka.
  • Developed Jenkins pipelines for continuous integration and deployment purpose.
  • Experience in working on analyzing snowflake datasets performance.
  • Worked on building pipelines using snowflake for extensive data aggregations.

Environment: Kafka, Spark, Sqoop, Hive, Azure, Databricks, Grafana, Jenkins, Azure Data Lake, Azure SQL, Jenkins, Grafana, Python, Shell, Microservices, Restful API's

We'd love your feedback!