Software System Engineer / Big Data Developer Resume
Manassas, VA
OBJECTIVE:
Big Data lead engineer, application developer or lead/mentor/manager position in cloud and big data environment that will effectively utilize my educational and professional experience to make positive contribution toward organizational goals and objectives, and position that will allow me to grow and learn. Always hungry to experiment with new technologies and methodologies that can enable better solutions.
SUMMARY:
- Highly motivated and curiosity driven Big Data and AWS developer with 15 years of experience in Software Development, management and Business Intelligence and over 5 years in Hadoop.
- Currently working on a ‘Global Data Warehouse’ project in semiconductor Company with over 10,000 employees and over 1TB of sensor data per day with focus on data collection, transformation and discoveries.
- Primarily focused on Hadoop (Spark, M/R, Hive, Sqoop), DB(Oracle, Teradata, MySQL, Mongo, Hbase, RedShift) and programming with Python, Java, Scala, R, Perl and PHP. Ability to develop end - to-end applications: from DB installation and performance tuning of huge volume of data, Hadoop data transformation and efficient storing to Web interface development.
- Ability to leverage statistical science and sense of exploration to mine large sets of structured and unstructured data.
- 15 years of experience in computer systems development and programming, over 5 years of big data development in Hadoop environment, 3 years of Cloud Computing, 9 years of semiconductor process control management with data analysis.
- Proven success in forming new teams, project management and scheduling. Ability to communicate efficiently.
- 4 semesters experience in teaching Master level course in Cloud Computing (Virtual Machines, AWS and Big Data) and at the local University in Virginia.
- Spent significant time in researching over 20 different databases (SQL and NoSQL) and performed many optimization experiments and data remodeling for finding best solutions and improving data usage for huge data volumes.
- Very good understanding of automation process, root cause analysis and process improvements.
- Experience in programming languages, Hadoop, Databases, distributed computer systems, math and statistic.
- Strong customer focus, providing good solutions and always more than expected. Ability to develop trust and help build teamwork. Love to learn about new technologies, research, teach and apply knowledge.
TECHNICAL SKILLS:
Relational databases: Oracle 10 & 11.2, Teradata, RedShift, MySQL (extensive experience with several different engines - InnoDB, MyISAM, Memory, CSV, Archive, Inforbright), MSSQL Server, Jethrodata (worked closely with original developers for new functionalities and improving performance), PostgreSQL, MemSQL, SQLite, Microsoft Access, Hive and Impala (extensive research in both)
NoSQL Databases: HBase, MongoDB, CouchDB, Cassandra, ElasticSearch and Solr
Tools: Oracle SQL Developer, Teradata Studio, MySQL Workbench, Heidi, Squirrel, Repl clients
Languages: Python (2.7, 3.5), JavaScript,SQL, Perl, JMP, Scala, Java, PHP (5.6 and 7.0), GO (1.6), Tcl, CSS, XML
Hadoop: Spark, Hive, Impala, Pig, Sqoop, M/R, Hbase, Yarn, Flume, Kafka, Zookeeper, Hue, Ambari, HDFS
Web: PHP, Node, JavaScript (JQuery, PivotTable, NodeJS, D3, AngularJS)
Web servers: Apache, NGINX, NodeJS, explored Caddy for HTTP2 and Security support.
Package Managers: YUM, Apt-get, Node NPM, GIT, Chocolatey, Python PIP, Perl PPM, Maven
Operating Systems: Experienced in Linux (RHEL, CentOS and Ubuntu) installing developer software and databases, troubleshooting, bash scripting, memory/cpu management and distributed systems. Windows Server . Win 7,10, MAC OS
Data Collection and Analysis: Capable of efficiently extracting and storing data in different formats from/to databases, HDFS, NAS and Linux file system. Experiences with Web HTTP and Rest APIs. Experience designing and architecting complex data flows using Big Data processing techniques and tools. Analysis via Spark, Python and R.
PROFESSIONAL EXPERIENCE:
Confidential, Manassas, VA
Software System Engineer / Big Data Developer
Responsibilities:
- On a daily basis developing different applications using Python, Java and Scala that utilize various data and construct ETL flows within Hadoop distributed ecosystem in 3 clusters of 40, 120 and 150 RHEL nodes in different parts of the world (Singapore, Virginia, Idaho). Document everything. Designing microservices.
- Typical ETL flow consist of moving data from Oracle (11.2), Teradata, MSSQL to Hadoop (HDFS, Hive, Hbase) using Sqoop (1.4.5) and custom scripts, performing data transformations before/after storing data utilizing Spark (1.6), Pig (0.14), Hive and custom Python/Perl scripts. Constructing flows from Hadoop to DB systems (MySQL 5.6, Postgress 9.2, HBase and Oracle). Coordination/scheduling via YARN managed clusters with Puppet config management.
- Cleaning, sorting and partitioning data. Storing using Parque, RCFile, Avro and Text file formats.
- Creating custom temporary and permanent UDF functions for dealing with specific cases with data for usage in Hive, Pig, Impala and Spark in HDFS and memory.
- Monitoring hadoop cluster system performance, behavior and spark/hive job execution trough Ambari, Hue, Web interface ports 50030-50070 and in Linux shell.
- Analyzing and data using JMP statistical software, Python and R for finding differences in groups, producing best fitting models for over 1 million charts based on long trend analysis and hourly alarms on out of population variations.
- Parallel processing with GOLang (1.6) for simultaneous distributed data access to carefully partitioned database systems (custom engines in MySql, Hbase) for fast data access and aggregation based on dynamic user requests.
- Creating interactive Web sites with PHP and AJAX calls accessing data directly from Impala via Thrift interface.
- Developing custom applications for accessing data from Hive in JMP and Excel using ODBC.
- Use Kerberos security authentication protocol and suggested solutions for custom groups.
- Developing Web RShiny apps in R with connecting, RDD and data frames manipulation in Spark.
- Exploring new solutions for improvement or replacement of existing technologies (Elastic search, Kibana, Cassandra, Docker). Part of the dev team for POC using Kafka and Flume integration.
- Visual representation in the browser using D3 libraries on Canvas and SVG and by using specialized tools (Tableau, HighCharts, DataNavigator) with user interaction using PHP and JQuery/Angular JS.
- International experience in forming, group mentoring and leading new Big Data teams in US (Boise 2013) and Asia (Singapore 2014, Taiwan 2015 and Japan 2015). Held training for new Big Data members, worldwide Hadoop discussions and architectural design improvements.
Software System Developer
Confidential
Responsibilities:
- Create queries for Oracle and MySQL database on a daily basis with integration with front end user interface by using C#, Perl, Visual Basic and JMP.
- Establish process control mechanisms, equipment monitoring methods, excursion monitoring, corrective actions and drive continuous improvement actions/plans for automation system development.
- Create interactive Web sites for different teams in the company for various purposes.
- Build engines for parsing data, sending notifications and automating that process on a daily basis.
Automation Software System Engineer
Confidential
Responsibilities:
- Used programming knowledge to integrate new tools into the IS system and make sure communication between systems meet standard quality.
- Create Web sites for IS teams and interactive interfaces for accessing data from different IS systems.
Confidential, VA
Cloud Computing adjunct Professor
Responsibilities:
- Teaching Master level course in Cloud Computing (AWS and fundamentals of Big Data) for several semesters to 20+ students at the time.
- Classes were mostly at the physical location in Richmond, VA and few times online teaching.
- Great experience and good feedbacks.
- Paused this activity due to forming Big Data teams in Asia .
Confidential
Lead Database and Web Administrator
Responsibilities:
- Worked in a world-wide construction company with over 1,000 employees as a Web/DB administrator leading a team of 5 people.
- Daily duties involved project planning for new Web site functions, following up on customer issues, planning for disaster recovery and report progress.
Database and Web Administrator
Confidential
Responsibilities:
- Supporting a system network of 200 computers, creating databases and web sites for different projects.
Confidential
Full time Database and Web Administrator
Responsibilities:
- Developed Web sites full time, programmed in JavaScript, PHP, Perl, Python and used MySQL for needed projects on Confidential Linux system as a platform.
- Requirements gathering, developing custom scripts and helping out senior developers.
- Creating experimental Web sites, programmed in PERL, PHP with MySQL for needed projects and used Confidential Linux system as a platform.
