We provide IT Staff Augmentation Services!

Cloud Big Data Engineer Resume

3.00/5 (Submit Your Rating)

New, JerseY

SUMMARY:

  • Senior software engineer with extensive experience in big data process, analysis, aggregation, migration, and ETL on Cloudera and AWS; system design and development for network planning; database application development; web application development; and AI application development.
  • Over 10 years in Java, and SQL development.
  • Strong Python and Tableau experience for data analysis and chart/dashboard creation.
  • SQL Query, View, and Stored Procedure
  • Optimization with Genetic Algorithms
  • AI System Development
  • Backend Business Logic and Database Development
  • Requirements Analysis and Design
  • Multitier Architecture Design and Development
  • Test Automation of Software Application Systems
  • System Performance Analysis and Enhancement
  • Lead Engineer for Various Software Development

SKILLS & COMPETENCIES

Databricks: Apache Spark, clusters and jobs, notebooks, Spark SQL, DataFrames and Datasets, Delta Lake, Spark MLlib, etc.

AWS: EC2, S3, Session Manager, CloudWatch, Lambda, SageMaker, Glue ETL Job, Crawler, Data Catalog, Athena, EMR, Presto, Redshift, KMS Key, IAM Roles, Policies, Security Configuration, Secret Manager, VPC, Security Group, etc.

Cloudera and Hadoop: Mapreduce, Spark, HDFS, Impala, Hive, HQL, Hive Metadata, Parquet, Hue, Workflow, Sqoop, etc.

Private Cloud: RHEL 7.5 Server, Kerberos, JupyterHub, Spark Edge Node, RStudio, VDI, etc.

Python: Dask, Dask Distribution cluster, PySpark, Pandas, DataFrame, Boto3, Multi - processing, Dash Visualization, etc.

Java: OOP - Encapsulation, Inheritance, and Polymorphismmulti-threading, JDBC, etc.

Other Programming Languages: C#, C++, C, VB, VBA, etc.

Database Systems: MySQL, PostgreSQL, Oracle, DENODO, SQL Server, Sybase, Access, etc.

Web Development: Spring Framework, ASP.NET, JavaScript, JSON, XML, HTML, etc.

Operating Systems: MacOS, Linux, Windows, Sun Solaris, etc.

Software Development Environments: GitHub, Jenkins, Spring-tool-suite, Gradle, Eclipse, Tableau, RapidMiner, OPNET, MS Visual Studio, SVN, Jira, etc.

PROFESSIONAL EMPLOYMENT:

Confidential, New Jersey

Cloud Big Data Engineer

Responsibilities:

  • Develop HQL queries for data analysis and data aggregation on Hive and Impala of Cloudera with Hue
  • Develop Spark Jobs to automate data aggregation and migrate data to AWS S3 and/or EMR
  • Develop ETL programs and job workflows to transfer data to AWS Redshift
  • ScholarOne Manuscripts Web Service Data Collection
  • Requirements Analysis and Design
  • Web Service Client Development
  • Redshift data ingestion
  • Develop Restful Web Service Client to collect ScholarOne Manuscripts data
  • Design Redshift data table schema
  • Develop Redshift data ingestion program
  • Data Lake Migration

Confidential

Cloud Big Data Engineer

Responsibilities:

  • Develop Python grogram and jobs to transfer data tables to AWS S3 and Redshift
  • Setup and configure AWS S3 and Glue environment
  • System test, troubleshooting, operation

Confidential, New York City

Big Data Engineer

Responsibilities:

  • Configurat Databrick Workers and create mapping folder to AWS S3 bucket
  • ETL big dataset to Databricks Delta Lake
  • Generate huge datasets and evaluate the performance of data process and analysis
  • Develop and test Databricks technologies on Databricks System
  • Evaluate the usability, stability, and performance of Databricks features
  • Make Evaluation documentation
  • Verint Audio file extraction and Lifecycle Management
  • Verint System analysis with C#
  • Big data collection with Python multi-process programming
  • AWS Lambda function implementation for Lifecycle Management
  • Application analysis and design
  • Python program and AWS Lambda function implementation
  • System test, troubleshooting, operation
  • Library Development for big data Analysis
  • Provide a Python library to export large dataset from Hive DB to HDFS
  • Provide a Python library to import large dataset from HDFS to Dask dataframe
  • Research and build a Dask distribution cluster to support large data set process and analysis
  • Application research and design
  • Library development, test, and troubleshooting
  • Library documentation and deployment
  • Cloud Server Performance Monitor

Confidential

Cloud Big Data Engineer

Responsibilities:

  • Automatically grab and collect CPU, memory, disk space usage info of all process running on the server
  • Automatically analysis every process usages and send an alarm to the owner of an idle process which is occupying large system resource
  • Automatically send an alarm to system administrator if system resource will exhaust soon
  • Performance Monitor design and development
  • Cron Job setup
  • Technical support for process owners

Confidential, New Jersey

Senior Software Engineer

Responsibilities:

  • IP-Optical Integration Design and Optimization System
  • Worked closely with domain experts to catch domain expert knowledge, implemented intelligent evolution methodology, and obtained optimization solution for the minimization of the network CapEx cost
  • Requirement gathering and analysis, Design, Implementation, Testing, Maintenance
  • Java object oriented programming (OOP): Encapsulation, Inheritance, and Polymorphism, etc.
  • Network Consolidation Design and Optimization System

We'd love your feedback!